Skip to main content
ITW Journal

ChatGPT can create 3D models in Blender

A few months ago, we might have said, “AI is getting so powerful that one day it could even create 3D models in programs like Blender or Cinema 4D.” That day is here.

Author
Angelo

Date
8 September 2026

Categories
ChatGPT, Astra, Codex, Blender


A beautifully made ad for Astra, the new ChatGPT model released a few days ago, shows what it can do: write copy, create content for brands and ad campaigns, and handle everyday tasks like ordering food or booking a tennis court.

But you can probably guess which part caught my attention. Astra can draw a 2D image, turn it into a 3D model, build a browser game around it, and even export the model for 3D printing.

Tools that generate 3D models from prompts or photos already exist, but the workflow in the ad is fascinating.


With a few days entirely to myself, I started exploring and testing this workflow to understand how it works, assess the results, and see how they might be useful in practice.


How do you model with ChatGPT?

If you use ChatGPT, you may have started with the smartphone app for everyday questions, then moved to a computer, using a browser like Chrome, Safari or Firefox.

There is also a ChatGPT desktop app that you can install like any other program. I had never felt the need to install it: I had always thought of ChatGPT as an online tool. The difference is that the desktop app can control other applications, and even interact with the operating system and much of your computer, with your permission, of course.

To have ChatGPT create 3D models on your computer, start by installing the desktop app: https://chatgpt.com/it-IT/download/

In both the OpenAI ad and the user videos shared on social media, Blender seems to be the most popular choice for these tests: https://www.blender.org/download/


Blender is free and open source software for 3D modeling, rendering and animation. It is popular with independent artists, though less common in architecture than 3ds Max or Cinema 4D. It has a stronger presence in animation, games, VFX and art.


Once both apps are installed, connect ChatGPT to Blender using MCP for Blender, a software component that links Codex (the mode in the ChatGPT desktop app that can carry out complex tasks, including controlling your computer) to Blender.

This lets ChatGPT read the scene, create objects, apply materials and perform advanced modeling tasks.

The setup takes some care, but it is straightforward. If you are unsure, ask ChatGPT to walk you through it one step at a time. If an instruction sounds too technical, ask it to explain in simpler terms.


Once connected, you can simply ask ChatGPT to create a 3D model of an object, specify the level of detail, and add reference images. Here is the result of my very first test.

Create a simplified 3D model of the original Nintendo Game Boy.

I deliberately started with a very basic prompt to see what the baseline result would look like. The first model is not bad, though many details differ from the original. The colors are quite accurate.


Next, I needed to work out how to improve the result, especially the details and the accuracy of the shapes. I started by asking ChatGPT why its Game Boy model looked so different from the original.

The Game Boy model looks fairly good, but some details still differ from the original. What instructions and reference images do you need to create a model that matches the reference as closely as possible? More generally, how can I get more accurate results when I ask you to model objects in Blender?

The result depends heavily on the amount and quality of the information ChatGPT receives before it starts modeling. In the first test, it had to find its own information about the Game Boy and piece it together, which explains the rough result.

For the second attempt, I found reference images myself and included them in the prompt.

  • 2D views

    Front, back, left, right, top and bottom views, looking straight at each surface

    Purpose: establish the overall shape and placement of details
  • 3D perspectives

    Angled views from both sides, including the back

    Purpose: understand depth, curves and transitions
  • Close-ups

    Close-ups of buttons, ports, dials, screws, seams and lettering

    Purpose: capture and reproduce fine details
  • Dimensions

    Height, width and depth; component sizes and spacing

    Purpose: keep proportions accurate
  • Text and logos

    Sharp, front-facing images, or vector files if available

    Purpose: reproduce typography and layout accurately
  • Hidden parts

    Images of the object opened up, its internal parts, or a version with a transparent shell

    Purpose: avoid guesswork

Your instructions should make four things clear:

Which reference takes priority

For example: “Use the front photo for proportions and the side photos only for depth.”

How accurate the model needs to be

Should it look right in a render, or be accurate enough to manufacture? Manufacturing requires verified measurements and tolerances.

Which details matter most

The outline, corner radii, recesses, lettering, internal parts, or wear.

How to handle unseen areas

Ask for another reference, or flag the area as an approximation.

Here is an example of a more detailed prompt:

Create an accurate, full-scale replica of [object, model and version] in Blender.

Use [image-name.jpg] as the main reference. The overall dimensions are in [image-name.jpg]. Prioritize the outline, curves, and the position and depth of details.

Start with the geometry and a neutral material. Overlay orthographic views on the references and correct any differences before adding materials and lettering.

Do not invent details that the references do not show. List any missing information and identify the parts that remain approximate.

Keep components separate and editable. All variations and renders must come from the updated Blender model.

Deliver the .blend file, comparison renders, and a short list of remaining approximations.
Copy

There is no fixed maximum or ideal number of images, but ChatGPT suggests 12 to 20 images with clear filenames such as “front.png”, “right_side.png” and “button_detail.png”. What matters is that each photo adds information: a few sharp, complementary views are more useful than lots of blurry, near-identical images.


Here is the second attempt, using five reference images and a slightly more structured prompt.

It is still not an exact match, and there is room for more detail, but it already looks much closer to the real Game Boy.


Staying in the same chat, I then moved on to image generation and asked ChatGPT to transform this image:

Give the Game Boy in this image a clear glass shell. Change the background from gray to black. Keep the details and image format. Add internal electronic components visible through the glass. Turn on the red LED to the left of the screen and show Super Mario on the display.
Copy

Using the same approach, I asked ChatGPT to model the DeLorean from Back to the Future as accurately as possible. I added around ten reference drawings and photos found through Google.

Create an accurate replica of the DeLorean from Back to the Future in Blender. Use the "blueprint" image as the main reference. Prioritize the outline, curves, and the position and depth of details. Start with the geometry and a neutral material.

Overlay orthographic views on the references and correct any differences before adding materials and lettering. Do not invent details that the references do not show. List any missing information and identify the parts that remain approximate.

Keep components separate and editable. All variations and renders must come from the updated Blender model. Deliver the .blend file, comparison renders, and a short list of remaining approximations.
Copy

Here is the result:

The process takes just a few minutes. Again, the model is quite basic.

Careful, non-destructive modeling, attention to detail, clean topology, sensible polygon counts, bevels and imperfections all help produce a highly realistic final image. Accuracy and detail also reflect the care and professionalism of the artist.

From my tests so far, these models do not have enough detail to hold up in a close-up. They do, however, make excellent starting points for images where detail can be added through AI image generation.



Turn this image into a realistic photo of the DeLorean from Back to the Future. Refine the simplified shapes and add realistic details. Keep the aspect ratio and camera framing. Place the DeLorean on a New York street at night, with traffic and people moving around it. Copy

With a 3D model, you can create as many starting images as you need in seconds. You can show the object from any angle, while the object itself stays the same, keeping the final images consistent. Small details may vary, but the overall geometry stays close to the model.

I used Kling to turn each image into a short animation, then edited them together in After Effects, with a little color correction to adjust the contrast and colors.

I wanted to try an object with simple shapes that could still make striking images. The Tesla Cybertruck came to mind. I followed the same workflow: I chose a camera angle in the Blender viewport, then asked ChatGPT to place the vehicle in the landscape shown in a reference photo.

For now, the main strength of this workflow is simple: in just a few minutes, you can create a 3D model and capture it from any angle, then turn those views into a series of striking, consistent images.

Astra can handle complex tasks, but the quality of the result still depends on our choices. Model accuracy has its limits, yet there is already plenty here to explore an idea and see where it leads.

The next step is to give Astra a real project. How far can we push the accuracy? And how much of this workflow can become part of our everyday work? ■

ANGELO FERRETTI

I’m Angelo. For over 20 years, my two biggest passions have been teaching and creating images for architecture and interior design.


italiano · english