๐“๐”€๐“ฎ๐“ผ๐“ธ๐“ถ๐“ฎ ๐“ฃ๐“ฎ๐”๐“ฝ๐Ÿ“-๐“ฝ๐“ธ-๐“˜๐“ถ๐“ช๐“ฐ๐“ฎ๐ŸŒ‡

Snapshot 2026-08-03 23:56:39 UTC ยท version 1

โ— published
INDEXABLE MARKDOWN SNAPSHOT

Research document

Open canonical .md

๐“๐”€๐“ฎ๐“ผ๐“ธ๐“ถ๐“ฎ ๐“ฃ๐“ฎ๐”๐“ฝ๐Ÿ“-๐“ฝ๐“ธ-๐“˜๐“ถ๐“ช๐“ฐ๐“ฎ๐ŸŒ‡

๐“ ๐“ฌ๐“ธ๐“ต๐“ต๐“ฎ๐“ฌ๐“ฝ๐“ฒ๐“ธ๐“ท ๐“ธ๐“ฏ ๐“ป๐“ฎ๐“ผ๐“ธ๐“พ๐“ป๐“ฌ๐“ฎ๐“ผ ๐“ธ๐“ท ๐“ฝ๐“ฎ๐”๐“ฝ-๐“ฝ๐“ธ-๐“ฒ๐“ถ๐“ช๐“ฐ๐“ฎ ๐“ผ๐”‚๐“ท๐“ฝ๐“ฑ๐“ฎ๐“ผ๐“ฒ๐“ผ/๐“ถ๐“ช๐“ท๐“ฒ๐“น๐“พ๐“ต๐“ช๐“ฝ๐“ฒ๐“ธ๐“ท ๐“ฝ๐“ช๐“ผ๐“ด๐“ผ.

โญ Citation

If you find this paper and repo helpful for your research, please cite it below:


@inproceedings{zhou2023vision+,
  title={Vision+ Language Applications: A Survey},
  author={Zhou, Yutong and Shimada, Nobutaka},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
  pages={826--842},
  year={2023}
}

๐ŸŽ‘ News

[!TIP] Version 1.0 (All-in-one version) can be found here and will be stop updating from 24/02/29.

  • [24/02/29] Update "Awesome Text to Image" Version 2.0! Paper With Code and Other Related Works will also be gradually updated in March.
  • [23/05/26] ๐Ÿ”ฅ Add our survey paper "Vision + Language Applications: A Survey" and a special Best Collection list!
  • [23/04/04] "Vision + Language Applications: A Survey" was accepted by CVPRW2023.
  • [20/10/13] Awesome-Text-to-Image repo is created.

To Do

Content

Description

  • In the last few decades, the fields of Computer Vision (CV) and Natural Language Processing (NLP) have been made several major technological breakthroughs in deep learning research. Recently, researchers interested in combining semantic information and visual information in these traditionally independent fields. A number of studies have been conducted on text-to-image synthesis techniques that transfer input textual descriptions (keywords or sentences) into realistic images.

  • Papers, codes, and datasets for the text-to-image task are available here.

๐ŸŒ Markdown Format:

Paper With Code

  • Text to Face๐Ÿ‘จ๐Ÿป๐Ÿง’๐Ÿ‘ง๐Ÿผ๐Ÿง“๐Ÿฝ
    • (ECCV 2024) PreciseControl: Enhancing Text-To-Image Diffusion Models with Fine-Grained Attribute Control, Rishubh Parihar et al. [Paper] [Project]
    • (arXiv preprint 2024) [๐Ÿ’ฌ Dataset] 15M Multimodal Facial Image-Text Dataset, Dawei Dai et al. [Paper]
    • (arXiv preprint 2024) [๐Ÿ’ฌ 3D] Portrait3D: Text-Guided High-Quality 3D Portrait Generation Using Pyramid Representation and GANs Prior, Yiqian Wu et al. [Paper]
    • (CVPR 2024) CosmicMan: A Text-to-Image Foundation Model for Humans, Shikai Li et al. [Paper] [Project]
    • (ICML 2024) Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric Regularization, Jinlu Zhang et al. [Paper] [Code]
    • (NeurIPS 2023) Inserting Anybody in Diffusion Models via Celeb Basis, Ge Yuan et al. [Paper] [Project]
    • (IJACSA 2023) Mukh-Oboyob: Stable Diffusion and BanglaBERT enhanced Bangla Text-to-Face Synthesis, Aloke Kumar Saha et al. [Paper] [Code]
    • (SIGGRAPH 2023) [๐Ÿ’ฌ 3D] DreamFace: Progressive Generation of Animatable 3D Faces under Text Guidance, Longwen Zhang et al. [Paper] [Project] [HuggingFace]
    • (CVPR 2023) [๐Ÿ’ฌ 3D] High-Fidelity 3D Face Generation from Natural Language Descriptions, Menghua Wu et al. [Paper] [Code] [Project]
    • (CVPR 2023) Collaborative Diffusion for Multi-Modal Face Generation and Editing, Ziqi Huang et al. [Paper] [Code] [Project]
    • (Pattern Recognition 2023) Where you edit is what you get: Text-guided image editing with region-based attention, Changming Xiao et al. [Paper] [Code]
    • (arXiv preprint 2022) Bridging CLIP and StyleGAN through Latent Alignment for Image Editing, Wanfeng Zheng et al. [Paper]
    • (ACMMM 2022) Learning Dynamic Prior Knowledge for Text-to-Face Pixel Synthesis, Jun Peng et al. [Paper]
    • (ACMMM 2022) Towards Open-Ended Text-to-Face Generation, Combination and Manipulation, Jun Peng et al. [Paper]
    • (BMVC 2022) clip2latent: Text driven sampling of a pre-trained StyleGAN using denoising diffusion and CLIP, Justin N. M. Pinkney et al. [Paper] [Code]
    • (arXiv preprint 2022) ManiCLIP: Multi-Attribute Face Manipulation from Text, Hao Wang et al. [Paper]
    • (arXiv preprint 2022) Generated Faces in the Wild: Quantitative Comparison of Stable Diffusion, Midjourney and DALL-E 2, Ali Borji, [Paper] [Code] [Data]
    • (arXiv preprint 2022) Text-Free Learning of a Natural Language Interface for Pretrained Face Generators, Xiaodan Du et al. [Paper] [Code]
    • (Knowledge-Based Systems-2022) CMAFGAN: A Cross-Modal Attention Fusion based Generative Adversarial Network for attribute word-to-face synthesis, Xiaodong Luo et al. [Paper]
    • (Neural Networks-2022) DualG-GAN, a Dual-channel Generator based Generative Adversarial Network for text-to-face synthesis, Xiaodong Luo et al. [Paper]
    • (arXiv preprint 2022) Text-to-Face Generation with StyleGAN2, D. M. A. Ayanthi et al. [Paper]
    • (CVPR 2022) StyleT2I: Toward Compositional and High-Fidelity Text-to-Image Synthesis, Zhiheng Li et al. [Paper] [Code]
    • (arXiv preprint 2022) StyleT2F: Generating Human Faces from Textual Description Using StyleGAN2, Mohamed Shawky Sabae et al. [Paper] [Code]
    • (CVPR 2022) AnyFace: Free-style Text-to-Face Synthesis and Manipulation, Jianxin Sun et al. [Paper]
    • (IEEE Transactions on Network Science and Engineering-2022) TextFace: Text-to-Style Mapping based Face Generation and Manipulation, Xianxu Hou et al. [Paper]
    • (CVPR 2021) TediGAN: Text-Guided Diverse Image Generation and Manipulation, Weihao Xia et al. [Paper] [Extended Version][Code] [Dataset] [Colab] [Video]
    • (FG 2021) Generative Adversarial Network for Text-to-Face Synthesis and Manipulation with Pretrained BERT Model, Yutong Zhou et al. [Paper]
    • (ACMMM 2021) Multi-caption Text-to-Face Synthesis: Dataset and Algorithm, Jianxin Sun et al. [Paper] [Code]
    • (ACMMM 2021) Generative Adversarial Network for Text-to-Face Synthesis and Manipulation, Yutong Zhou. [Paper]
    • (WACV 2021) Faces a la Carte: Text-to-Face Generation via Attribute Disentanglement, Tianren Wang et al. [Paper]
    • (arXiv preprint 2019) FTGAN: A Fully-trained Generative Adversarial Networks for Text to Face Generation, Xiang Chen et al. [Paper]

<๐ŸŽฏBack to Top>

  • Specific Issues๐Ÿค”
    • (arXiv preprint 2026) [๐Ÿ–ผ๏ธ Aesthetic Dataset] Moonworks Lunara Aesthetic Dataset, Yan Wang et al. [Paper] [Dataset]
    • (arXiv preprint 2026) [๐Ÿ“ธ Variation Dataset] Moonworks Lunara Aesthetic II: An Image Variation Dataset Yan Wang et al. [Paper] [Dataset]
    • (arXiv preprint 2025) [๐Ÿ’ฌ Differentiable Object Counting] YOLO-Count: Differentiable Object Counting for Text-to-Image Generation, Guanning Zeng et al. [Paper]
    • (arXiv preprint 2024) [๐Ÿ’ฌ Gender Bias Alignment] PopAlign: Population-Level Alignment for Fair Text-to-Image Generation, Shufan Li et al. [Paper] [Code]
    • (arXiv preprint 2024) [๐Ÿ’ฌ Fine-Grained Feedback] Beyond Thumbs Up/Down: Untangling Challenges of Fine-Grained Feedback for Text-to-Image Generation, Katherine M. Collins et al. [Paper]
    • (CVPR 2024-Best Paper) [๐Ÿ’ฌ Human Feedback] Rich Human Feedback for Text-to-Image Generation, Youwei Liang et al. [Paper]
    • (ICLR 2024) [๐Ÿ’ฌ Unauthorized Data] DIAGNOSIS: Detecting Unauthorized Data Usages in Text-to-image Diffusion Models, Zhenting Wang et al. [Paper] [Code]
    • (CVPR 2024) [๐Ÿ’ฌ Open-set Bias Detection] OpenBias: Open-set Bias Detection in Text-to-Image Generative Models, Moreno D'Incร  et al. [Paper]
    • (arXiv preprint 2024) [๐Ÿ’ฌ Spatial Consistency] Getting it Right: Improving Spatial Consistency in Text-to-Image Models, Agneet Chatterjee et al. [Paper] [Project] [Code] [Dataset]
    • (arXiv preprint 2024) [๐Ÿ’ฌ Safety] SafeGen: Mitigating Unsafe Content Generation in Text-to-Image Models, Xinfeng Li et al. [Paper] [Code]
    • (arXiv preprint 2024) [๐Ÿ’ฌ Aesthetic] Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation, Daiqing Li et al. [Paper] [Project] [HuggingFace]
    • (EMNLP 2023) [๐Ÿ’ฌ Text Visualness] Learning the Visualness of Text Using Large Vision-Language Models, Gaurav Verma et al. [Paper] [Project]
    • (arXiv preprint 2023) [๐Ÿ’ฌ Against Malicious Adaptation] IMMA: Immunizing text-to-image Models against Malicious Adaptation, Yijia Zheng et al. [Paper] [Project]
    • (arXiv preprint 2023) [๐Ÿ’ฌ Principled Recaptioning] A Picture is Worth a Thousand Words: Principled Recaptioning Improves Image Generation, Eyal Segalis et al. [Paper]
    • โญโญ(NeurIPS 2023) [๐Ÿ’ฌ Holistic Evaluation] Holistic Evaluation of Text-To-Image Models, Tony Lee et al. [Paper] [Code] [Project]
    • (ICCV 2023) [๐Ÿ’ฌ Safety] Rickrolling the Artist: Injecting Backdoors into Text Encoders for Text-to-Image Synthesis, Lukas Struppek et al. [Paper] [Code]
    • (arXiv preprint 2023) [๐Ÿ’ฌ Natural Attack Capability] Intriguing Properties of Diffusion Models: A Large-Scale Dataset for Evaluating Natural Attack Capability in Text-to-Image Generative Models, Takami Sato et al. [Paper]
    • (ACL 2023) [๐Ÿ’ฌ Bias] A Multi-dimensional study on Bias in Vision-Language models, Gabriele Ruggeri et al. [Paper]
    • (FAACT 2023) [๐Ÿ’ฌ Demographic Stereotypes] Easily Accessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale, Federico Bianchi et al. [Paper]
    • (arXiv preprint 2023) [๐Ÿ’ฌ Robustness] Evaluating the Robustness of Text-to-image Diffusion Models against Real-world Attacks, Hongcheng Gao et al. [Paper]
    • (CVPR 2023) [๐Ÿ’ฌ Adversarial Robustness Analysis] RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation With Natural Prompts, Han Liu et al. [Paper]
    • (arXiv preprint 2023) [๐Ÿ’ฌ Textual Inversion] Is This Loss Informative? Speeding Up Textual Inversion with Deterministic Objective Evaluation, Anton Voronov et al. [Paper] [Code]
    • (arXiv preprint 2022) [๐Ÿ’ฌ Interpretable Intervention] Not Just Pretty Pictures: Text-to-Image Generators Enable Interpretable Interventions for Robust Representations, Jianhao Yuan et al. [Paper]
    • (arXiv preprint 2022) [๐Ÿ’ฌ Ethical Image Manipulation] Judge, Localize, and Edit: Ensuring Visual Commonsense Morality for Text-to-Image Generation, Seongbeom Park et al. [Paper]
    • (arXiv preprint 2022) [๐Ÿ’ฌ Creativity Transfer] Inversion-Based Creativity Transfer with Diffusion Models, Yuxin Zhang et al. [Paper]
    • (arXiv preprint 2022) [๐Ÿ’ฌ Ambiguity] Is the Elephant Flying? Resolving Ambiguities in Text-to-Image Generative Models, Ninareh Mehrabi et al. [Paper]
    • (arXiv preprint 2022) [๐Ÿ’ฌ Racial Politics] A Sign That Spells: DALL-E 2, Invisual Images and The Racial Politics of Feature Space, Fabian Offert et al. [Paper]
    • (arXiv preprint 2022) [๐Ÿ’ฌ Privacy Analysis] Membership Inference Attacks Against Text-to-image Generation Models, Yixin Wu et al. [Paper]
    • (arXiv preprint 2022) [๐Ÿ’ฌ Authenticity Evaluation for Fake Images] DE-FAKE: Detection and Attribution of Fake Images Generated by Text-to-Image Diffusion Models, Zeyang Sha et al. [Paper]
    • (arXiv preprint 2022) [๐Ÿ’ฌ Cultural Bias] The Biased Artist: Exploiting Cultural Biases via Homoglyphs in Text-Guided Image Generation Models, Lukas Struppek et al. [Paper]

<๐ŸŽฏBack to Top>

  • 2025

    • (arXiv preprint 2025) GenExam: A Multidisciplinary Text-to-Image Exam, Zhaokai Wang et al. [Paper]
    • (arXiv preprint 2025) RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation, Aviv Slobodkin et al. [Paper]
    • (arXiv preprint 2025) An Empirical Study of GPT-4o Image Generation Capabilities, Sixiang Chen et al. [Paper]
  • 2024

    • (arXiv preprint 2024) Flow Generator Matching, Zemin Huang et al. [Paper]
    • (EMNLP 2024) Kandinsky 3: Text-to-Image Synthesis for Multifunctional Generative Framework, Vladimir Arkhipkin et al. [Paper] [Code] [Project]
    • (arXiv preprint 2024) Data Extrapolation for Text-to-image Generation on Small Datasets, Senmao Ye and Fei Liu [Paper]
    • โญโญ(arXiv preprint 2024) Imagen 3, ImagenTeam-Google [Paper]
    • (arXiv preprint 2024) MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis, Wanggui He et al. [Paper]
    • (Kuaishou) Kolors: Effective Training of Diffusion Model for Photorealistic Text-to-Image Synthesis, Sixian Zhang et al. [Paper] [Code] [Project]
    • (CVPR 2024) [๐Ÿ’ฌHuman Preferences] Learning Multi-dimensional Human Preference for Text-to-Image Generation, Sixian Zhang et al. [Paper] [Code] [Project]
    • (CVPR 2024) [๐Ÿ’ฌ Text-to-layout โ†’ Text+Layout-to-Image] Grounded Text-to-Image Synthesis with Attention Refocusing, Quynh Phung et al. [Paper] [Project] [Code]
    • (arXiv preprint 2024) Dimba: Transformer-Mamba Diffusion Models, Zhengcong Fei et al. [Paper]
    • (arXiv preprint 2024) [๐Ÿ’ฌ Generation and Editing] MultiEdits: Simultaneous Multi-Aspect Editing with Text-to-Image Diffusion Models, Mingzhen Huang et al. [Paper] [Project]
    • (arXiv preprint 2024) AutoStudio: Crafting Consistent Subjects in Multi-turn Interactive Image Generation, Junhao Cheng et al. [Paper] [Project] [Code]
    • (arXiv preprint 2024) TheaterGen: Character Management with LLM for Consistent Multi-turn Image Generation, Junhao Cheng et al. [Paper] [Project] [Code]
    • (CVPR 2024) Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following, Yutong Feng et al. [Paper] [Project] [Code]
    • (arXiv preprint 2024) CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching, Dongzhi Jiang et al. [Paper] [Project] [Code]
    • (arXiv preprint 2024) TextCraftor: Your Text Encoder Can be Image Quality Controller, Yanyu Li et al. [Paper]
    • (CVPR 2024) ECLIPSE: A Resource-Efficient Text-to-Image Prior for Image Generations, Maitreya Patel et al. [Paper] [Project] [Code] [Hugging Face]
    • (arXiv preprint 2024) SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data, Jialu Li et al. [Paper] [Project] [Code]
    • (ICLR 2024) PixArt-ฮฑ: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis, Junsong Chen et al. [Paper] [Project] [Code] [Hugging Face]
    • (arXiv preprint 2024) PixArt-ฮฃ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation, Junsong Chen et al. [Paper]
    • (arXiv preprint 2024) PIXART-ฮด: Fast and Controllable Image Generation with Latent Consistency Models, Junsong Chen et al. [Paper]
    • (CVPR 2024) Discriminative Probing and Tuning for Text-to-Image Generation, Leigang Qu et al. [Paper] [Project]
    • (CVPR 2024) RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization, Mengqi Huang et al. [Paper] [Project]
    • โญ(arXiv preprint 2024) SDXL-Lightning: Progressive Adversarial Diffusion Distillation, Shanchuan Lin et al. [Paper] [HuggingFace] [Demo]
    • โญ(arXiv preprint 2024) RealCompo: Dynamic Equilibrium between Realism and Compositionality Improves Text-to-Image Diffusion Models, Xinchen Zhang et al. [Paper] [Code]
    • (arXiv preprint 2024) Learning Continuous 3D Words for Text-to-Image Generation, Ta-Ying Cheng et al. [Paper] [Project] [Code]
    • (arXiv preprint 2024) DiffusionGPT: LLM-Driven Text-to-Image Generation System, Jie Qin et al. [Paper] [Project] [Code]
    • (arXiv preprint 2024) DressCode: Autoregressively Sewing and Generating Garments from Text Guidance, Kai He et al. [Paper] [Project]

<๐ŸŽฏBack to Top>

  • 2023
    • (arXiv preprint 2023) CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation, Zineng Tang et al. [Paper] [Project] [Code]
    • (arXiv preprint 2023) DiffBlender: Scalable and Composable Multimodal Text-to-Image Diffusion Models, Sungnyun Kim et al. [Paper] [Code] [Project]
    • (arXiv preprint 2023) ElasticDiffusion: Training-free Arbitrary Size Image Generation, Moayed Haji-Ali et al. [Paper] [Project] [Code] [Demo]
    • (ICCV 2023) BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion, Jinheng Xie et al. [Paper] [Code]
    • (arXiv preprint 2023) Late-Constraint Diffusion Guidance for Controllable Image Synthesis, Chang Liu et al. [Paper] [Code]
    • (arXiv preprint 2023) An Image is Worth Multiple Words: Multi-attribute Inversion for Constrained Text-to-Image Synthesis, Aishwarya Agarwal et al. [Paper]
    • โญ(arXiv preprint 2023) UFOGen: You Forward Once Large Scale Text-to-Image Generation via Diffusion GANs, Yanwu Xu et al. [Paper]
    • (ICCV 2023) ITI-GEN: Inclusive Text-to-Image Generation, Cheng Zhang et al. [Paper] [Code] [Project]
    • (arXiv preprint 2023) Mini-DALLE3: Interactive Text to Image by Prompting Large Language Models, Zeqiang Lai et al. [Paper] [Code] [Demo] [Project]
    • (arXiv preprint 2023) [๐Ÿ’ฌEvaluation] GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment, Dhruba Ghosh et al. [Paper] [Code]
    • โญ(arXiv preprint 2023) Kandinsky: an Improved Text-to-Image Synthesis with Image Prior and Latent Diffusion, Anton Razzhigaev et al. [Paper] [Code] [Demo] [Demo Video] [Hugging Face]
    • โญโญ(ICCV 2023) Adding Conditional Control to Text-to-Image Diffusion Models, Lvmin Zhang et al. [Paper] [Code]
    • (ICCV 2023) DiffCloth: Diffusion Based Garment Synthesis and Manipulation via Structural Cross-modal Semantic Alignment, Xujie Zhang et al. [Paper]
    • (ICCV 2023) Unsupervised Compositional Concepts Discovery with Text-to-Image Generative Models, Nan Liu et al. [Paper] [Code] [Project]
    • (arXiv preprint 2023) Text-to-Image Generation for Abstract Concepts, Jiayi Liao et al. [Paper]
    • (arXiv preprint 2023) T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation, Kaiyi Huang et al. [Paper] [Code] [Project]
    • (arXiv preprint 2023) [๐Ÿ’ฌ Evaluation] Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis, Xiaoshi Wu et al. [Paper] [Code]
    • (arXiv preprint 2023) Towards Unified Text-based Person Retrieval: A Large-scale Multi-Attribute and Language Search Benchmark, Shuyu Yang et al. [Paper] [Code] [Project]
    • (arXiv preprint 2023) Synthesizing Artistic Cinemagraphs from Text, Aniruddha Mahapatra et al. [Paper] [Code] [Project]
    • (arXiv preprint 2023) Detector Guidance for Multi-Object Text-to-Image Generation, Luping Liu et al. [Paper]
    • (arXiv preprint 2023) A-STAR: Test-time Attention Segregation and Retention for Text-to-image Synthesis, Aishwarya Agarwal et al. [Paper]
    • (arXiv preprint 2023) [๐Ÿ’ฌEvaluation] ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models, Maitreya Patel et al. [Paper] [Code] [Project]
    • โญ(arXiv preprint 2023) StyleDrop: Text-to-Image Generation in Any Style, Kihyuk Sohn et al. [Paper] [Project]
    • โญโญ(arXiv preprint 2023) Prompt-Free Diffusion: Taking "Text" out of Text-to-Image Diffusion Models, Xingqian Xu et al. [Paper] [Code] [Hugging Face]
    • โญโญ (SIGGRAPH 2023) Blended Latent Diffusion, Omri Avrahami et al. [Paper] [Code] [Project]
    • (CVPR 2023) [๐Ÿ’ฌControllable] SpaText: Spatio-Textual Representation for Controllable Image Generation, Omri Avrahami et al. [Paper] [Project]
    • โญโญ (arXiv 2023) The Chosen One: Consistent Characters in Text-to-Image Diffusion Models, Omri Avrahami et al. [Paper] [Code] [Project]
    • (CVPR 2023) [๐Ÿ’ฌStable Diffusion with Brain] High-resolution image reconstruction with latent diffusion models from human brain activity, Yu Takagi et al. [Paper]
    • (arXiv preprint 2023) BLIP-Diffusion: Pre-trained Subject Representation for Controllable Text-to-Image Generation and Editing, Dongxu Li et al. [Paper]
    • (arXiv preprint 2023) [๐Ÿ’ฌEvaluation] LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation, Yujie Lu et al. [Paper] [Code]
    • (arXiv preprint 2023) P+ : Extended Textual Conditioning in Text-to-Image Generation, Andrey Voynov et al. [Paper] [Project]
    • (arXiv preprint 2023) Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models, Xuhui Jia et al. [Paper]
    • (ICML 2023) TR0N: Translator Networks for 0-Shot Plug-and-Play Conditional Generation, Zhaoyan Liu et al. [Paper] [Code] [Hugging Face]
    • (ICLR 2023) [๐Ÿ’ฌ3D]DreamFusion: Text-to-3D using 2D Diffusion, Ben Poole et al. [Paper (arXiv)] [Paper (OpenReview)] [Project] [Short Read]

This HTML preview is truncated for page performance. The canonical Markdown file contains the complete snapshot.

MARKDOWN METRICS
10369words
9headings
933links
1code blocks
MDRSS ASSESSMENT
Evidence46/100high confidence
Why MDRSS assigned this score
  • Production catalog audit 2026-08-04
  • Taxonomy classified from title, annotation, source and Markdown signals
  • Agent usefulness evaluated from structure, procedures, examples, evidence and retrieval value
Evidence (1)

Discussion 0

Sign in to join the discussion.