ChatGPT’s image generation model gets an amazing upgrade for businesses

OpenAI has given image generation within ChatGPT a huge upgrade, making the AI much more capable of delivering legible text in images and more.

The move sees OpenAI switch to an entirely new model, replacing the DALL-E model that OpenAI has used until now and which delivered inferior results compared to dedicated AI image services such as Midjourney.

OpenAI claims image generation is now a “primary capability” of its language models, and has thus built an advanced image generator into GPT-4o, which is currently the default language model used for ChatGPT queries.

The company claims the new model “excels at accurately rendering text, precisely following prompts, and leveraging 4oโ€™s inherent knowledge base and chat context – including transforming uploaded images or using them as visual inspiration”.

Accurate text in AI images

The new model heralds a massive leap forward in AI’s ability to accurately render text. Most previous models have struggled to legibly deliver more than two or three words, with spelling errors or blurred text common.

The GPT-4o image generator appears to have cracked the problem, capable of rendering large sections of text accurately.

For example, I asked ChatGPT to “generate an image of a teacher writing the opening lines of Macbeth on a whiteboard” and it delivered the following image:

ChatGPT image showing teacher at whiteboard

When I asked it to create a cover of PC Pro magazine, with the lead headline “Windows 12 – Everything You Need To Know”, it produced the following:

ChatGPT image generation magazine cover

It generated the rest of the text on the cover itself, choosing coverlines that you might well find on a tech magazine… and using the magazine’s logo without permission.

For businesses, this could make ChatGPT much more useful for generating promotional images to be used on social media or even print promotions, such as below.

ChatGPT image for business to promote a sale

Creating characters

Another potential business use for ChatGPT’s new image generator is the creation of characters that the AI can remember from prompt to prompt.

Previously, if you generated a character or image of a person, then asked ChatGPT to use that same character/person in a subsequent image, the best you could hope for would be a faint likeness. Now it can maintain the character’s appearance across multiple generations, allowing businesses to create characters that can be used in multiple advertising campaigns.

To give a very crude example, I generated this image of a cartoon puppy:

ChatGPT image generation of a puppy

And then asked the AI to use that character in a sofa store promotion, where the character is sitting on a sofa wearing fluffy slippers…

ChatGPT image of a puppy with slippers on

It works for photo-realistic characters too, so you could generate the equivalent of Howard from the Halifax ads or the Go Compare singer for your own company, without the considerable expense of retaining models/actors.

The new image model is live now and available to all account tiers, including free, although there are limits on the number of images you can generate on free accounts. Developers will get access to the model via API “in the next few weeks”, according to OpenAI.

Whether it solves OpenAI’s struggles to understand idioms remains to be seen…

Avatar photo
Barry Collins

Barry has 25 years of experience working on national newspapers, websites and magazines. He was editor of PC Pro and is co-editor and co-owner of BigTechQuestion.com. He has published a number of articles on TechFinitive covering data, innovation and cybersecurity.