Home / Case studies / AI Level Design That Builds Walkable Levels for Game Studios
Case study · Gaming · Generative AI & agentic AI
AI Level Design That Builds Walkable Levels for Game Studios
A US game development studio wanted AI level design to replace the handmade graybox blockouts its designers assembled block after block for every title. Brainy Neurals built a prompt-to-level pipeline that uses small AI agents to turn one written brief into a 3D scene. Designers type the brief in plain words, then open the finished level file in the engine the team already uses. Every object in the level is now placed & sized by the pipeline, so a layout revision takes one prompt edit.
One prompt
Starts every new level
In the engine
Where the team walks it
Prompt edit
How a layout gets revised
Published October 2026
At a glance
What problem did this solve?
Building graybox 3D levels meant placing every object by hand before anyone could walk the space. Iteration was slow, & each new idea cost the same manual effort.
What did Brainy Neurals build?
Brainy Neurals built a multi-agent AI pipeline for a US game development studio. The pipeline turns a written prompt into a structured 3D scene & a walkable graybox file.
What changed after it went live?
A level idea now goes from text to a walkable graybox without anyone placing blocks. Designers spend their time judging layouts instead of building them, & a revision takes one prompt edit.
Who else could use this?
Prompt-to-3D generation fits any team that plans a space before building it. Architecture firms & warehouse planners sketch layouts the way game studios do, & so do training-simulation builders.
| Industry | Sub-vertical | Client | Engagement | Timeline | Capabilities | Delivery |
|---|---|---|---|---|---|---|
| Gaming | 3D level design tooling | US game development studio | Prompt-to-level generation pipeline | Not disclosed | Generative AI & agentic AI | Project-based delivery |
Why was level blockout so slow?
AI level design had to fix one slow stage at a US game development studio that builds 3D titles level by level. Every level starts as a graybox, a set of plain blocks a player can walk before any art exists. Blockout decides whether a layout earns its art pass or gets cut from the game.
The studio wanted to test far more layout ideas each cycle than its designers could block out by hand. Until then, generative AI development had only produced images & text for the team, with no space anyone could walk.
- Each block was dragged into place by hand
- A revision meant rebuilding the whole layout
- Every pitch waited on a walkable build
- Each new variant cost the same hours again
A 2013 survey of procedural content generation found that hand-made game content no longer keeps pace with rising production costs[1]. The pipeline in this case study was built to close that gap for one studio.
Before the pipeline, each layout lived on paper until someone built it in the editor, block by block.
What do studios usually reach for?
Studios usually pick one of four routes, & each of them makes sense somewhere.
| Approach | What it gets right | Where it stops | Who it still suits |
|---|---|---|---|
| Hand blockout | Full control over every block | Hours per layout & per revision | Hero levels & final passes |
| Rule-based generation | Endless variation, fast | Can’t follow a written brief | Roguelike runs & terrain |
| Learned generators | Pick up a game’s style | Need existing levels to learn from | Games with levels to spare |
| Prompt-to-level pipeline, our route | A brief becomes a walkable file | Schema & validation work up front | Teams testing many spatial ideas |
Research on machine-learned generation names the catch with the third route[2].
Those models train on existing levels, so a working library has to exist before they can generate anything.
How we built the AI level design pipeline
Brainy Neurals built the pipeline as a chain of small agents, the central AI agent development decision on this build. A planning agent reads each brief & splits the work between specialist agents. Each specialist then owns a single concern, such as the layout or the rules that tie objects together.
Their answers merge into one structured scene. One strict schema defines exactly what that scene may contain.
The model decides what goes where, & deterministic code turns that decision into geometry.
Our first decision was what the model would be allowed to produce. We ruled out free-form output early, because geometry code can’t negotiate with prose. Every agent answers inside the schema, & any other answer counts as a failed attempt.
Our second decision was where generation stops & code takes over. The model never emits a vertex or a mesh. It only describes the level, & an open-source mesh library builds real geometry from that description.
Every agent answers to the same contract, & that shared contract is what holds the whole chain together.
The model decides what goes where, & deterministic code turns that decision into geometry.
Language stops at the schema line, & only numbers that pass validation ever become geometry.
The technology stack we used
Every layer in this stack exists to keep language apart from geometry. Our technology selection questions were the usual ones about what the model decides & what code guarantees. The list stays short on purpose, because fewer moving parts leave fewer places for a scene to go wrong.
| Layer | What we used | Why | What we ruled out |
|---|---|---|---|
| Language model | A frontier model family | Read spatial briefs more reliably than the others in our tests | Two other families we benchmarked |
| Orchestration | One agent per concern | Small, focused prompts hold detail | A single do-everything prompt |
| Scene schema | One strict schema per scene | Geometry code needs guarantees, not prose | Free-form model output |
| Validation | Schema & spatial checks | Bad scenes bounce back before they become geometry | Trusting the first answer |
| Mesh generation | An open-source mesh library | Deterministic geometry from validated numbers | Asking the model for meshes |
| Output | One standard 3D file | Opens in the engines the team already uses | Engine-locked exports |
How does one prompt become a level?
One prompt becomes a level in six steps, in the exact order the pipeline runs them.
A scene that fails a check goes straight back for repair before any geometry gets built.
- A designer writes the level in plain language, covering its spaces, paths, obstacles & overall mood.
- A planning agent reads the brief & splits it into separate jobs for the specialist agents downstream.
- Specialist agents turn each job into entries in one structured scene, with a set position & size for every object.
- A validation pass checks the scene against the schema & the spatial rules, then sends failures back for repair.
- The mesh stage reads the validated scene & builds real 3D geometry from it, one primitive for each entry.
- The pipeline writes one standard 3D file, & the team opens it in the engine to walk the level.
A person appears only at the two ends of that six-step chain.
What broke in the first builds?
Three problems broke the first builds, & they showed up in this order.
The first scenes read well but parsed badly. A field would go missing or a type would change, & some lists arrived as prose. The geometry code then crashed on input it had every right to trust.
The second failure was spatial rather than structural. Scenes began to pass validation & still made no sense, because objects overlapped or floated above the floor. Some walls sealed off the only path.
A scene could be valid JSON & still describe a level nobody could play.
The third problem was overload inside a single prompt. That one prompt handled the whole brief at once & dropped requirements as briefs grew longer. The model was stretched too thin for the job it had.
Research on language-guided 3D scene generation treats object positioning as a problem of its own[3]. The published fix gives a solver explicit spatial constraints instead of trusting a model’s raw coordinates. Weeks like these are when clients bring in specialist engineers instead of learning every lesson the slow way.
How we got past each one
We fixed each of the three problems separately, & every fix took several rounds to find.
Structure
We locked every agent to the schema & made every answer validate on arrival. A failed check goes back with the exact error, so the agent can repair its own output. Nothing malformed has reached the geometry code since that change landed.
Space
We added a spatial pass behind the schema pass. That pass checks for overlapping or floating objects & for spaces a player can’t reach. Each violation comes back with its coordinates, & a scene becomes geometry only after it clears both gates.
Split
We broke the single big prompt into the agent chain the pipeline runs today. Each agent holds one concern, so detail stopped falling out of long briefs.
Whitebox means the same thing as graybox, incidentally, & no one agrees on which name came first.
Surfacing this kind of work before anything gets promised is what an AI proof of concept is for.
Every generated scene passes the schema gate & then the spatial gate, or it goes back for repair.
What changed after go-live?
Going live changed how every layout at the studio gets built, from the first idea to each revision.
| What | Before | After |
|---|---|---|
| Getting to a walkable layout | Hand placement in the editor | One written prompt |
| Objects placed by hand | Every object in the scene | None |
| Revising a layout | Rebuild the blockout | Edit the prompt & run again |
| Trying variants | One at a time, by hand | One prompt each |
| What a designer does | Builds the space | Judges the space |
We haven’t published a generation-time or hours-saved figure, because this build measured neither. We only print numbers we have actually measured on a build.
Day to day, AI level design changes where a designer’s attention goes. A layout idea now gets typed out instead of dragged into place. Every variant costs only one prompt, so exploring more ideas no longer means more building for the team.
The graybox itself stayed the same, & only the way it gets made has changed.
Planning spaces by hand before you build?
Tell us what your team still blocks out by hand, from game levels to warehouse racking. We’ll tell you where a written prompt could draft that layout instead.
What is running today
The AI level design pipeline is in use at the studio today, on the same agent chain the first version proved. A designer writes the brief & starts the run. The finished file then opens in the engine the team already uses. The schema is still the contract between language & geometry, & nothing becomes a mesh without passing it.
Every level the pipeline produces still starts as one prompt, just as it did on the first working run. Judging layouts is now the real design work, & typing the brief is the whole build step.
In the running pipeline, a failed scene goes back for repair before the level ever assembles.
What would we do differently?
Four changes would have saved us time on this build, & we now start every similar project with them.
Design the schema before the prompts
We wrote the prompts first & had to rework them once the schema landed. The schema turned out to be the product itself.
Put a check at every handoff
Our first validator sat at the very end, so one early mistake surfaced three stages late. Errors got cheap once checks ran at every handoff.
Walk every candidate before judging it
A generated level nobody can walk through is only a picture of a design. Early reviews leaned on top-down views, & clean layouts read wrong at player height. Now every candidate gets walked before anyone judges it.
Keep the model away from numbers
Letting the model fix coordinates is tempting when a scene is close. Every nudge invites drift, so repairs go back through the schema like everything else.
A generated level nobody can walk through is only a picture of a design.
Where else does prompt-to-3D fit?
Prompt-to-3D generation turns a written description into a walkable 3D layout, & it fits wherever space gets planned before it gets built.
| Industry | The equivalent problem | What changes in the build |
|---|---|---|
| Construction | Clients want walkable massing before drawings exist | Real dimensions & code-driven constraint rules |
| Logistics | warehouse rack layouts sketched before steel is ordered | Racking sizes & aisle-width rules in the schema |
| Retail | Store planners test fixture layouts for each site | A fixture vocabulary & footfall constraints |
| Training simulation | Safety teams need many scenario spaces at low cost | Hazard placement rules & scenario variants |
| Virtual production | Scenes blocked out before any set is built | Camera-aware layouts & per-shot variants |
Porting the pipeline takes a new object vocabulary in the schema plus the constraint rules of that domain, with a validation pass tuned to both.
The same prompt-to-3D pattern can plan a warehouse, with aisles & racking blocked out before anything is built.
Questions buyers usually ask
How it works
Can AI generate a 3D level from a text prompt?
Yes, as long as the model describes the level & code builds the geometry. A language model turns the prompt into a structured scene, & a mesh stage turns that scene into geometry. Asking a model for raw 3D geometry directly doesn’t hold up.
What is a graybox level, & why start there?
A graybox is a level built from plain untextured blocks to test layout & flow before any art exists. Teams start there because moving a gray block is cheap & moving finished art is expensive.
How is this different from procedural generation?
Classic procedural generation follows fixed rules or noise, so it can’t read a designer’s brief. A prompt-driven pipeline starts from written intent, & a designer who describes a level gets that level instead of a random one.
Time, cost & rollout
How long does a prompt-to-level pipeline take to build?
A working proof of concept on your own content usually takes a few weeks. The scene schema & the agent split each need a dedicated pass, & so do the validation rules. A production pipeline follows once generated levels start surviving review unedited.
How much does an AI level design build cost?
Cost depends on the variety of objects & the constraint rules, plus how much content you already have. Brainy Neurals scopes the build after a short conversation & a close look at your content, then quotes a fixed price. An AI readiness assessment tells you first whether your briefs & assets can drive one.
Does AI level generation replace level designers?
No, it replaces the placement labor & leaves the judgment about what plays well to your designers. They still decide what makes a space worth playing, & the pipeline gets them to that decision sooner.
What should a prompt build for you?
Levels, layouts, massing, racking & scenario spaces can all start as a prompt. If your team plans it before building it, tell us about the project & a written brief could probably draft it.
Services behind this case study
Brainy Neurals built this case study from the six services below.
Generative AI development
Text-to-3D pipelines built to return structured output instead of prose.
AI agent development
Multi-agent orchestration where specialist agents split a task & check each other’s work.
RAG development
Generation grounded in your own design documents & asset catalogs
Document AI
Messy briefs & design documents turned into structured data your tools can read.
Hire AI developers
Generative AI engineers who extend your team for the weeks when the schema fights back.
AI in construction
Layout & massing generation for construction, where space gets planned before it gets built.
An AI proof of concept is the fastest way to test this on your own content. AI consulting helps you choose the model family & the line between language & geometry first. An AI readiness assessment tells you up front whether your briefs can drive a schema. The industries hub shows where else this generation pattern already runs.
Similar case studies
Brainy Neurals shipped each of these three builds into a real environment instead of a slide deck.
Overhead Line Geometry Measurement
Stereo cameras on a moving train measure wire geometry, with inference running on the train.
AI Diet Assistant for Gastroenterology
Clinical dietary guidance generated under review gates, live in a healthcare setting.
Personalised AI Meal Planning for Chronic Care
Structured meal plans built from messy personal health data, grounded & reviewable.
Cite this case study
Trivedi, Viral & Patel, Mitesh. AI Level Design That Builds Walkable Levels for Game Studios. Brainy Neurals, August 2026. https://brainyneurals.com/case-studies/ai-3d-level-design/
Sources cited on this page
- Hendrikx M, Meijer S, Van der Velden J, Iosup A. Procedural Content Generation for Games: A Survey. ACM Transactions on Multimedia Computing, Communications, and Applications. 2013;9(1):1-22. DOI 10.1145/2422956.2422957.
- Summerville A, Snodgrass S, Guzdial M, Holmgård C, Hoover AK, Isaksen A, Nealen A, Togelius J. Procedural Content Generation via Machine Learning (PCGML). IEEE Transactions on Games. 2018;10(3):257-270. DOI 10.1109/TG.2018.2846639.
- Yang Y, Sun FY, Weihs L, VanderBilt E, Herrasti A, Han W, Wu J, Haber N, Krishna R, Liu L, Callison-Burch C, Yatskar M, Kembhavi A, Clark C. Holodeck: Language Guided Generation of 3D Embodied AI Environments. CVPR 2024. arXiv 2312.09067.








