ChatGPT’s new Images 2.0 model is surprisingly good at generating text
Identifying whether an image was created by AI or a human is becoming increasingly difficult, thanks to the advances in OpenAI’s ChatGPT Images 2.0 model. Just two years ago, AI struggled with simple tasks like generating restaurant menus, often introducing awkwardly misspelled dishes. Now, the latest model can produce accurate, polished images containing readable text that would fit seamlessly into real-world settings.
Technological Improvements
Previous AI image generators, such as DALL-E 3, were notorious for spelling errors due to their reliance on diffusion models that reconstructed images from noise and often neglected fine details like text. Researchers responded by exploring autoregressive models, which operate similarly to large language models and show more promise in handling text within images. While OpenAI hasn't revealed which approach powers Images 2.0, the improvements are clear.
Key Features of ChatGPT Images 2.0
- Text Rendering: The model can generate clear, accurate text, including complex elements like menus, comic panels, and user interface components.
- Multilingual Support: ChatGPT Images 2.0 now better handles non-Latin scripts, including Japanese, Korean, Hindi, and Bengali.
- Creative Flexibility: It can produce images in a variety of styles and sizes, and generate multiple images from a single prompt.
- Knowledge Cutoff: Model knowledge is current up to December 2025, which may limit accuracy on very recent topics.
- Efficiency: While not instant, the model generates even complex scenes within a few minutes.
Availability and Usage
The new model is accessible to all ChatGPT and Codex users, with more advanced features available to paid subscribers. Developers can also leverage the gpt-image-2 API, with pricing varying based on output quality and resolution.
To read the detailed original article by Amanda Silberling, visit TechCrunch.