A review of two text-to-image models that pushed the field forward in photo-realism, zero-shot ability, and the way text and images share a latent space. Co-authored with Spriha Joshi and Lucas Giovanni Uberti-Bona Marin, with thanks to Prof. Siamak Mehrkanoon.
Review of OpenAI's DALL-E 2 and DALL-E
Authors: Spriha Joshi, Lucas Giovanni Uberti-Bona Marin, Lucas-Andreï Thil.
Intro
Recent improvements in text-to-image representations enabled the creation of impressive image generation algorithms. In this report, we review two such models called DALL-E and DALL-E 2 that have achieved state-of-the-art results in terms of photo-realism, zero-shot capabilities and context relation between textual and image inputs. In the following sections we will describe their approach and highlight some key concepts that made this technology different from the rest. We will also describe key results and discuss its strengths and weaknesses. Lastly, some possible future works and risks that are presented by DALL-E will be presented.
Approach
Both DALL-E and DALL-E 2 use a two-stage architecture where the first stage consists of training a latent space representational algorithm between text and images, while stage 2 trains a decoder model reconstructing the images from that shared latent representation. DALL-E 2 shows significant improvements from DALL-E by using its CLIP ranking algorithm for the latent representation instead of dVAE, and uses the GLIDE diffusion model that was introduced in an intermediate paper in 2021 [9]. This was also done by improving the training data representation.
Stage 1 · VQ-VAE and dVAE
DALL-E 1 used an adaption of the Vector Quantized Variational Auto Encoder called Dynamic Variational Auto Encoder, or dVAE. The use of this method allows to reduce the parameter space necessary to encode images, similar to the problem solved by using the convolutional filters introduced in the course.
The differences with a classical VAE is that it uses a discrete latent space representation instead of a continuous one. This is motivated by several factors:
- Most of real world data favors a discrete representation.
- VAEs latent representation suffers from continuous approximations and need to use KL-divergence between true posterior and its approximated form of independent Gaussians. Information is lost in this process.
dVAE introduces a discrete codebook to generate the latent space, removing the need to use KL-divergence from classical VAEs. This removes the need of rounding off continuous data and increases the representational accuracy. Furthermore, during the training the errors is backpropagated via gradient descent directly to the prior's output and avoiding the latent space encoding contrary to classical VAEs where its latent space gets affected. The prior's training is initialized with a uniform distribution and updates during the training, while the posterior is deterministic in range [0,1] which is why the need of KL-divergence is removed.
Stage 1 · CLIP
Contrastive Language-Image Pre-training (CLIP) is a neural network model returning the best caption for a given image. It connects textual and visual information together by computing the cosine similarity of the two inputs which are fed at the same time during training. Tokens are generated in the latent space representing similar elements by grouping them close together, and increase the distance between non-related ones. CLIP achieved state of the art zero-shot capabilities in comparison experiments over various datasets [14]. The latent zi describes the aspects of the image representation by clip, xT its textual one and is encoded into the bipartite latent representation (zi,xT). The latent xT allows to reconstruct the image from the residual information by applying a DDIM [10] inversion to x using the decoder while conditioning on zi. This allows for various types of manipulations such as interpolating intermediate spaces in order to generate new images. By computing the difference vector of two different representational tokens, intermediate images can be generated by traversing CLIP's concept space between those. The produced images are related to the two initial ones by a progressive variation of the representational latent space tokens. Because the latent space represents visual and textual information together, it allows to do text interpolation to produce these intermediate images through a spherical interpolation between the CLIP image embedding before being decoded.
CLIP was used in DALL-E to perform ranking over the produced images, but was not to perform any internal model representation. This is not the case for DALL-E 2 where it is actively used in the model to learn the latent space representation before being frozen to move into the second stage of the model. The main idea was that because CLIP uses the same token representation for both text and images and achieving better results, it should be used as the representation model.
DALL-E 2 uses the initial CLIP algorithm by feeding a caption to produce an image instead of providing an image to generate a textual description of it. Because of this, DALL-E is also referred as 'unCLIP'.
Stage 2 · DALL-E
During the training stage in phase 2 DALL-E concatenates the BPE-encoded text captions with the image tokens generated dVAE using a token for separation between image and text. Then a decoder-only auto-regressive transformer is used to predict the image tokens. This presents an alternative to the Generative Adversarial Networks (GANs) introduced in the course for image generation.
The auto-regressive (AR) transformer [17] contains the following elements:
- A positional encoder: transformers don't take into consideration the internal order of a given sequence. To be able to encode the position of each token in a sequence positional encoding is used. Using positional encoding a sine and cosine value are alternatively added to even and odd position tokens embedding.
- A masked multi-head attention block: these blocks aim to learn interdependencies between the different elements of a sequence. The attention is computed by learning three matrices (Query, Key and Value) that are multiplied individually with the original sequence.
- Where M is a mask. A mask is a matrix containing minus-infinity and 0 values. The values above the diagonal have minus-infinity value to avoid using later tokens to predict those coming before. Similarly under the diagonal we can have different patterns of values that will force the attention to 0. This block is called multi-headed because different values can be learnt for Q, K and V and different masks (M) can be applied leading to different attention scores that may focus on separate possible relationships between tokens in a sequence.
These elements are sometimes part of residual blocks as the ones studied in the course which allow the architecture to self-configure the number needed to obtain the optimal representation.
The image tokens are then passed through the decoder of the dVAE to obtain the final images. The images will be ranked based on the similarity between images and captions found by CLIP [11] and the best ones will be displayed for the user.
Stage 2 · DALL-E 2
The first step of training for stage 2 of DALL-E 2 is learning a prior P(zi|y) that can produce CLIP image embedding from a given caption embedding. The second step is learning a decoder that can produce images from the image embedding and (possibly) the caption embedding.
The prior is learnt in two different ways:
- By using an AR model similar to that used in DALL-E.
- By using a diffusion model conditioned on the caption. The diffusion model [15] [4] learns the prior by: 1) progressively adding Gaussian noise to the sequence in discrete steps; 2) trying to revert the noise generation by making a Markov assumption that the current noisy sequence depends only on the previous step.
By using this approach the prior can be learnt effectively and guide the generation of image tokens.
Another diffusion model is again used to generate images from the given image tokens (as a decoder) by guiding the diffusion towards the desired tokens. And two other diffusion models are used to further upsample the image to a higher quality. These upsampling diffusion models use the dropout method introduced during the course to improve performance.
Results · DALL-E
To analyse DALL-E's performance, a comparison was made with AttnGAN, DM-GAN, and DF-GAN [18] [19] [16]. These were some of the state-of-the-art text-to-image generative models that existed prior to DALL-E.
For this analysis, two datasets, namely, MS-COCO and CUB-200 were used. Microsoft Common Objects in Context, or MS-COCO, is a captioning dataset containing 328 thousand images along with their natural language descriptors [6] whereas Caltech-UCSD Birds-200-2011 or CUB-200 is another dataset containing approximately 12 thousand images of only birds [3]. The dataset was extended by [14] by adding natural language descriptors to each image. These two datasets are benchmark datasets and have been used for evaluation of generative models.
The evaluation was done using three metrics:
- Inception Score (IS): a metric that asses the quality of the images generated based on two properties, namely, variety of images generated and the uniqueness of each image [7] [2].
- Fréchet Inception Distance (FID): a measure of similarity that also assesses the quality of images generated compared to the true image [5].
- Human Evaluation: for this, evaluators were shown a set of images and the caption when necessary. For each image, the human evaluators were asked to choose the output image that matched the caption the best or looked the most realistic or natural. Thus, three properties were evaluated, namely, caption matching, photorealism and sample diversity [12].
Before computing the FID and IS scores, a Gaussian Filter is applied to the images generated from AttnGAN, DM-GAN and DF-GAN. As mentioned before, DALL-E utilises dVAE, which is able to capture low frequency details well but compresses high frequency details making the final images produced by DALL-E slightly low in resolution. Hence, for a fair comparison of DALL-E's performance, by applying a Gaussian Filter, the images produced by other models are blurred with varying degrees of filter kernel sizes [13].
Figure 2 shows the FID and IS scores obtained by all models under evaluation as a function of the kernel size. On MS-COCO dataset, DALL-E performs the best by achieving a low FID and high IS score. However, on CUB-200 dataset, DALL-E performance drops drastically with a low IS and high FID score. The reason for this is attributed to the difference in distribution of images that DALL-E was trained on and the distribution of images in the CUB dataset which contains only images of birds.
Moreover, on human evaluation, the images produced by DALL-E on MS-COCO dataset received a majority vote on caption similarity 93% of the times and on photorealism 90% of the times. Thereby concluding that human evaluators preferred the images from DALL-E.
Moreover, Figure 1 gives an overview of the qualitative comparison of images generated with DALL-E vs images generated with other prior state-of-the-art models. Evidently, DALL-E is able to produce images of high quality that look natural and include well defined objects from the caption and are composed in a proper manner. On the other hand, images produced from the rest of the models struggle in producing well defined objects. The images look synthetic with overlaying patterns or textures and also seem distorted in some cases.
Lastly, an unexpected result of DALL-E was seen in its ability to perform Image-to-Image translation guided by natural language [13]. This can be seen from Figure 4 in Appendix A. DALL-E is able to perform operations such as edge detection flipping it on x-axis or y-axis, converting it to gray scale or changing the color of the object suggesting that it is able to perform object segmentation to a certain extent.
Results · DALL-E 2
DALL-E 2's performance was compared to GLIDE [9] on MS-COCO dataset using the three metrics introduced in the section before. The results also reflect performance of the two versions of DALL-E 2: with an auto-regressive prior or a diffusion prior that were discussed in the sections before.
Figure 3 summarises FID scores obtained by various text to image generation models. Among other Zero-Shot models, DALL-E 2 with a diffusion prior sets a new state-of-the-art record on the FID score of 10.39 on the MS-COCO dataset.
To gain insights into the CLIP Latent space, an experiment was conducted where images with typographical attacks were fed into the model. The attack is made by overlaying the text on the main object of the image. The aim of such attacks is to confuse the model into producing images of objects described in the text instead of the main original object that is present in the image. Two observations can be made using the results shown in Figure 5 in Appendix A. Firstly, the images produced by DALL-E 2 clearly show the apple that was subjected to the attack. Thus, showing DALL-E 2's ability to stir away from the attacks successfully. However, a second observation can be made which is that the text that is reproduced are random hinting on some limitations of DALL-E 2's encoding of texts. This is further discussed in the next section.
Lastly, a human evaluation of DALL-E 2 is also performed. To summarise, for photorealism, GLIDE is preferred over DALL-E 2 (or unCLIP) whereas for sample diversity, DALL-E 2 is shown to perform better than GLIDE. Moreover, among the two versions of DALL-E 2, the version with diffusion prior performs better than the version with an auto-regressive prior on all three criteria of human evaluation.
Strengths and Weaknesses
DALL-E has many capabilities as listed below:
- Given a caption, DALL-E is able to independently generate images of multiple objects, controlling their attributes as well as the relationship between. To some extent, it is able to even replicate one objects multiple times [1].
- It is able to interpret multiple interpretations of words, if they exist, and produce images depicting them [1].
- Its ability to contextualise both in terms of, the style of painting or drawing mentioned in the caption and getting the details, colors and textures right, and surprisingly also in terms of time [1]. Unexpectedly, it is also able to link text with certain styles from an era as can be seen in Figure 6 in Appendix A [8].
- Another very unique ability is that it can control the angle and location from which a scene is seen. An example of this is shown in Figure 8 in Appendix A [1].
It has many applications, mainly in the creative fields. For example, fashion design, interior design, product design, generation of emojis, stock images and many more.
However, it also comes with its own set of limitations which are listed below:
- The biggest limitation of DALL-E and DALL-E 2 is that it is extremely computationally expensive. It requires a large amount of computation power which only very few organisations have. Moreover, OpenAI has been very opaque with sharing exact details of the architecture and the datasets that have been used to build and train the model. This makes it very hard for a thorough analysis of the model. Nevertheless we present some observations below.
- As we also saw from the experimental results of DALL-E 2, it struggles with the encoding of textual information such as spellings, negations of words and some expressions like "similar" [8].
- Lastly, as the complexity of the caption increases in terms of the objects, their attributes and the associations among them, DALL-E is seen to struggle with compositionality and counting (the objects). It starts to confuse some attributes and the placement of the objects with respect to each other [8].
Future Work
As seen in Figure 2, DALL-E's performance drops drastically on the CUB-200. One possible way to improve this is by fine tuning DALL-E on datasets with specialised distributions such as CUB-200 [13]. A major risk with models like DALL-E is how it can be used for generation of harmful content or deceptive content like fake news. Therefore, a discussion needs to be held on how DALL-E should be released to the public, if at all. Another question that it raises is regarding who gets the copyrights over the images produced by models such as DALL-E. We can also think about the future improvements such as video applications or photo-realism. Those do not come free from social ramifications such as massively disrupting creative art industries by the quick and versatile generative abilities of DALL-E and similar methods.
References
- Jan 2021. OpenAI blog post.
- S. Barratt and R. Sharma. A note on the inception score. arXiv:1801.01973, Jun 2018.
- X. He and Y. Peng. Fine-grained visual-textual representation learning. IEEE Transactions on Circuits and Systems for Video Technology, 30(2):520 to 531, Feb 2020.
- J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. NeurIPS, 33:6840 to 6851, 2020.
- T. Kynkäänniemi, T. Karras, M. Aittala, T. Aila, and J. Lehtinen. The role of imagenet classes in Fréchet inception distance, 2022.
- T.-Y. Lin et al. Microsoft COCO: Common objects in context, 2014.
- D. Mack. A simple explanation of the inception score, Mar 2019.
- G. Marcus, E. Davis, and S. Aaronson. A very preliminary analysis of DALL-E 2, p. 14.
- A. Nichol et al. GLIDE: towards photorealistic image generation and editing with text-guided diffusion models, 2021.
- A. N. Prafulla Dhariwal. Diffusion models beat GANs on image synthesis, 2021.
- A. Radford et al. Learning Transferable Visual Models From Natural Language Supervision. arXiv:2103.00020, Feb. 2021.
- A. Ramesh et al. Hierarchical Text-Conditional Image Generation with CLIP Latents. arXiv:2204.06125, Apr. 2022.
- A. Ramesh et al. Zero-Shot Text-to-Image Generation. arXiv:2102.12092, Feb. 2021.
- S. Reed, Z. Akata, B. Schiele, and H. Lee. Learning deep representations of fine-grained visual descriptions, 2016.
- J. Sohl-Dickstein, E. A. Weiss, N. Maheswaranathan, and S. Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. arXiv:1503.03585, 2015.
- M. Tao et al. DF-GAN: a simple and effective baseline for text-to-image synthesis, 2020.
- A. Vaswani et al. Attention Is All You Need. arXiv:1706.03762, Dec. 2017.
- T. Xu et al. AttnGAN: fine-grained text to image generation with attentional GANs, 2017.
- M. Zhu, P. Pan, W. Chen, and Y. Yang. DM-GAN: dynamic memory GANs for text-to-image synthesis, 2019.
Appendix · Additional Figures
Fig. 1. Qualitative comparison of image outputs from DALL-E with AttnGAN, DM-GAN and DF-GAN.
Fig. 2. Inception Score and Fréchet Inception Distance as a function of the kernel size on MS-COCO (top) and on CUB-200 (bottom).
Fig. 3. Comparison of FID scores with various other text-to-image generators.
Fig. 4. Image-to-Image translation.
Fig. 5. Experimental results of images featuring typographical attacks.
Fig. 6. Caption: "Abraham Lincoln touches his toes while George Washington does chin-ups. Lincoln is barefoot. Washington is wearing boots."
Fig. 7. Caption: "A fish-eye lens view of an owl sitting in a field."
Fig. 8. Caption: "A red ball on top of a blue pyramid with the pyramid behind a car that is above a toaster."
All credits and rights of the images and contents respective to their creators, and great consideration towards their work.