I was as excited as any AI-aware engineer would be when I learned that OpenAI had announced a grand leap forward in the image generation capabilities of it’s 4o model. I’ve been using ChatGPT’s image-generation capabilities for a while, mostly to generate funny profile pictures for my privacy-conscious online friends who didn’t want to share their real-life face, but I needed a way to identify them in my contacts regardless. Interestingly, though there are many models out there, I’ve never felt a compulsion to try any other than DALL-E and (now) 4o. No Midjourney for me.
I read an interesting comment on Hacker News that describes the new image generation capabilities as “reasoning in pixel space” and used a Tic-Tac-Toe board as an example. Here is the exact text:
Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on.
To me, this prompt is a little boring. I am an avid chess player, though, so I immediately thought of a more fun application: could ChatGPT 4o’s new image generation capabilities reliably render either a 3D or a 2D-model of a chessboard in a particular position, much like the T.T.T. example above? Fortunately, chess benefits from algebraic notation–a textual format for representing positions–something Tic-Tac-Toe woefully lacks (and why, a curious mind asks). I knew I wouldn’t be able to feed it a 20-move game like the ones I normally play so I kept it simple: The London Opening.
Intro to London
London is considered a very basic opening, because White plays pretty much the exact same opening moves regardless of what Black does. Bearing in mind the historically limited capabilities of AI image generation, I considered it an ideal candidate due to the simplicity and shortness. To be explicit, this is the series of moves I was hoping to render with 4o:
1.d4 d5 2.Nf3 Nf6 3.Bf4
This is what this position should look like on a real (2D) board, courtesy of Chess.com:

I will leave the manifestation of a 3D board as an exercise to the reader. ChatGPT won’t be able to help you here, as I later found out.
First Attempt
Prompt 1
Here is the first prompt I used and the outcome:
Generate an image of a 3-dimensional chessboard displaying the London Opening, characterized by this algebraic notation:
1.d4 d5 2.Nf3 Nf6 3.Bf4
No humans need be present in the picture, only the chessboard in the CORRECT orientation, the pieces in the CORRECT position, and it can be sitting on a simple table with a simple background. The focus is on the board.
Result 1

For a zero-shot, and my admittedly less-than-stellar prompt (I was running a quick experiment, checkmate me) this isn’t bad. While you can tell there are some low-level issues like off-center pieces, incorrect rank/file labeling, and a really silly-looking knight on g1, the key problems are with the pieces that ARE on the board. I described the multitude of problems in a follow-up prompt.
Second Attempt
Prompt 2
In this prompt I used a continuance to explain what was wrong with the generated image, focusing on the extra knight, the missing pawns, the missing Queen, and the overall wrongness of the pictured position based on the input:
Close but not quite there. You got the style right, but the position is wrong in several ways:
1. There are three black knights on the board.
2. The black knight on f4 is two squares downfile of where it should be. I think the white bishop should be there.
3. There is no white queen on the board
4. I do not see the pawns in the center.
Try again
The purpose behind this was to provide slight corrections to the LLM model in the hopes that the second generation would be as high-quality as the first, but with the key improvements to bring it closer to the ideal final state. I had a feeling by this point already however that the task I was asking for was above the capabilities of the model in its current state. Still, I kept digging.
Result 2

Interestingly, our extra knight disappeared, but so too did other critical pieces: the white Queen is still missing, a white knight is now missing, and several pawns are missing from both sides of the board. But, the white bishop is a lot closer to the correct square than it was before, the only problem is it is currently sitting where the white knight should go. I do appreciate that the d2 pawn is missing, because it almost looks like the model moved the pawn out of the way to get the bishop to its intended square, which makes more sense than an image depicting a bishop that has “jumped” a blocking pawn, a move which is not legal in chess. However, it appears to have handled this issue by “disappearing” the pawn entirely. Gotcha!
Third Attempt
Prompt 3
It was at this point that I decided that my human-written prompts weren’t quite cutting it (simplistic as they were). So, I used a different 4o chat to create an AI-engineered prompt, as follows [note: ChatGPT fluff removed for brevity].
Better, but still not quite there.
I will try telling you the exact piece layout to see if that helps you.
Below I will express the board as an 8×8 matrix indexed from a1 to h8 (bottom-left to top-right from White’s perspective), using piece abbreviations:
• Capital = White, lowercase = Black
• P = Pawn
• N = Knight
• B = Bishop
• R = Rook
• Q = Queen
• K = King
• . = Empty square
Here’s the coordinate layout (from rank 8 to rank 1):
8 | r n b q k b n r
7 | p p p . p p p p
6 | . . . . . . . .
5 | . . . p . . . .
4 | . . . . B . . .
3 | . . . . P N . .
2 | P P P P . P P P
1 | R N B Q K . . R
a b c d e f g h
Now, please try again
Result 3

In this image, we can see that the model’s generation capabilities kind of…fall off of a cliff. The rank and file labels are way off, even mixing up letters and numbers on the right side of the board. Also, white’s kingside rook has been replaced with a pawn, and…actually…*stares closer*…there is no white king on the board. Yay! ChatGPT just invented a foolproof way to win a chess game as black! Just generate the board already in a checkmated position!
Sarcasm aside, this is likely due to the extended context of the chat, which has now gone through three rounds of generation. One of the biggest drawbacks of AI is that as the context window lengthens, the responses typically get worse or less detailed. An example of this is an AI-led project I did recently to convert recipes from audio transcriptions to HTML. Initially, the model was programmed to “always” include the amounts of each ingredient, and infer them from general cooking/baking knowledge if they were missing (transcription garbled). After converting 50+ recipes, all quantities were completely gone, though the ingredients were still present. A reminder in a new prompt was required to fix the issue.
However, the “reinforcement of the rules” trick (it probably has a stuffy official name, admittedly I don’t know it) doesn’t appear to work with image generation.
Final Attempt
I decided to give it one last shot before calling it a day, still (perhaps erroneously) believing that it was the prompt that was the problem, and not the model. Granted, this image generation is lightyears ahead of what we had before, where most text was either only partially legible or complete gibberish, as well as other obvious artifacts of computer generation like hands with too many fingers or other small details that are missed or fudged.
Prompt 4
This was another “LLM optimized” notation generated by 4o for 4o: I’m not convinced it is any better than any of the previous attempts.
No, some pieces are still not correct. Here is another way of expressing the notation in text format, please generate an image EXACTLY according to the below piece arrangement:
Rank 8: [Black Rook on a8] [Black Knight on b8] [Black Bishop on c8] [Black Queen on d8] [Black King on e8] [Black Bishop on f8] [Black Knight on g8] [Black Rook on h8]
Rank 7: [Black Pawn on a7] [Black Pawn on b7] [Black Pawn on c7] [Empty on d7] [Black Pawn on e7] [Black Pawn on f7] [Black Pawn on g7] [Black Pawn on h7]
Rank 6: [Empty on a6] [Empty on b6] [Empty on c6] [Empty on d6] [Empty on e6] [Empty on f6] [Empty on g6] [Empty on h6]
Rank 5: [Empty on a5] [Empty on b5] [Empty on c5] [Black Pawn on d5] [Empty on e5] [Empty on f5] [Empty on g5] [Empty on h5]
Rank 4: [Empty on a4] [Empty on b4] [Empty on c4] [White Pawn on d4] [White Bishop on e4] [Empty on f4] [Empty on g4] [Empty on h4]
Rank 3: [Empty on a3] [Empty on b3] [Empty on c3] [Empty on d3] [White Pawn on e3] [White Knight on f3] [Empty on g3] [Empty on h3]
Rank 2: [White Pawn on a2] [White Pawn on b2] [White Pawn on c2] [White Pawn on d2] [Empty on e2] [White Pawn on f2] [White Pawn on g2] [White Pawn on h2]
Rank 1: [White Rook on a1] [White Knight on b1] [White Bishop on c1] [White Queen on d1] [White King on e1] [Empty on f1] [Empty on g1] [White Rook on h1]
Result 4

I would write a paragraph with some analysis of the model’s output in this pass, but it seems unnecessary given that the generated image is so similar to the ones generated by previous prompts, but with even more mistakes. Too many to correct in a follow-up prompt without spoiling the original context chain, in my opinion. One additional note on the similarity of the images: it is so noticeable, I wondered if OpenAI may be using some sort of caching of graphical data behind-the-scenes to save on processing power.
After four passes, I think it’s reasonable to conclude that as of March 2025, when this article was written and published, 4o is not QUITE capable of rendering an image depicting a scenario as complicated as a chessboard in a particular position. As a follow-up experiment, I think it would be interesting to see if it could generate an accurate endgame position, with far fewer pieces on the board. Perhaps the reduced complexity, combined with those “LLM optimized” board layouts (what’s wrong with algebraic notation, ChatGPT??) some interesting results could develop.
Conclusion
I’ve been using AI and AI image generation for over a year now, like most techies. Obviously, it’s become mainstream in the ensuing time, the point where family members are reaching out to me to determine if I’ve heard of “Claude” or “Chat Gee Pee Tee”. The answer I give them is the same:
- I use these tools daily, and I use multiples of them for a variety of purposes.
- AI is the next technological revolution. Get on this train while you can.
- If you don’t know how to use AI, ask it. (Unless you’re a chessmaster, in which case I would wait until the next model update.)
One fun fact to close: Google’s Gemini is another model that is fairly good at image generation, but it has no problem generating images of copyrighted figures. Here’s one I made for my daughter, a professed Princess Peach aficionado:

Further Reading
- Generating Images with Google Gemini – Google Gemini (formerly Bard) docs
- Creating Images in ChatGPT – OpenAI help docs
- Starting today, GPT-4o is going to be incredibly good at image generation – reddit.com
- HN comment thread on the OpenAI announcement

