Home / Case studies / AI Level Design That Builds Walkable Levels for Game Studios

Case study · Gaming · Generative AI & agentic AI

AI Level Design That Builds Walkable Levels for Game Studios

A US game development studio wanted AI level design to replace the handmade graybox blockouts its designers assembled block after block for every title. Brainy Neurals built a prompt-to-level pipeline that uses small AI agents to turn one written brief into a 3D scene. Designers type the brief in plain words, then open the finished level file in the engine the team already uses. Every object in the level is now placed & sized by the pipeline, so a layout revision takes one prompt edit.

  • #AILevelDesign
  • #GenerativeAI
  • #MultiAgentAI
  • #GameDevelopment
  • #GrayboxLevels

One prompt

Starts every new level

In the engine

Where the team walks it

Prompt edit

How a layout gets revised

Mitesh

Published October 2026

At a glance

What problem did this solve?

Building graybox 3D levels meant placing every object by hand before anyone could walk the space. Iteration was slow, & each new idea cost the same manual effort.

What did Brainy Neurals build?

Brainy Neurals built a multi-agent AI pipeline for a US game development studio. The pipeline turns a written prompt into a structured 3D scene & a walkable graybox file.

What changed after it went live?

A level idea now goes from text to a walkable graybox without anyone placing blocks. Designers spend their time judging layouts instead of building them, & a revision takes one prompt edit.

Who else could use this?

Prompt-to-3D generation fits any team that plans a space before building it. Architecture firms & warehouse planners sketch layouts the way game studios do, & so do training-simulation builders.

Industry Sub-vertical Client Engagement Timeline Capabilities Delivery
Gaming 3D level design tooling US game development studio Prompt-to-level generation pipeline Not disclosed Generative AI & agentic AI Project-based delivery

Why was level blockout so slow?

AI level design had to fix one slow stage at a US game development studio that builds 3D titles level by level. Every level starts as a graybox, a set of plain blocks a player can walk before any art exists. Blockout decides whether a layout earns its art pass or gets cut from the game.

The studio wanted to test far more layout ideas each cycle than its designers could block out by hand. Until then, generative AI development had only produced images & text for the team, with no space anyone could walk.

  • Each block was dragged into place by hand
  • A revision meant rebuilding the whole layout
  • Every pitch waited on a walkable build
  • Each new variant cost the same hours again

A 2013 survey of procedural content generation found that hand-made game content no longer keeps pace with rising production costs[1]. The pipeline in this case study was built to close that gap for one studio.

Manual blockout, with gray blocks placed one at a time into an unfinished layout & wiped back to empty when a revision lands BLOCKOUT BY HAND WAITING TO BE PLACED …AND MORE REVISION · REBUILD THE BLOCKOUT PLACED THIS PASS STILL EMPTY ONE BLOCK AT A TIME EVERY REVISION STARTS THE LAYOUT AGAIN Manual blockout, with gray blocks placed one at a time, stacked for small screens BLOCKOUT BY HAND WAITING TO BE PLACED ONE BLOCK AT A TIME 3 PLACED · 3 STILL EMPTY REVISION · REBUILD THE BLOCKOUT EVERY REVISION STARTS AGAIN

Before the pipeline, each layout lived on paper until someone built it in the editor, block by block.

What do studios usually reach for?

Studios usually pick one of four routes, & each of them makes sense somewhere.

Approach What it gets right Where it stops Who it still suits
Hand blockout Full control over every block Hours per layout & per revision Hero levels & final passes
Rule-based generation Endless variation, fast Can’t follow a written brief Roguelike runs & terrain
Learned generators Pick up a game’s style Need existing levels to learn from Games with levels to spare
Prompt-to-level pipeline, our route A brief becomes a walkable file Schema & validation work up front Teams testing many spatial ideas

Research on machine-learned generation names the catch with the third route[2].
Those models train on existing levels, so a working library has to exist before they can generate anything.

How we built the AI level design pipeline

Brainy Neurals built the pipeline as a chain of small agents, the central AI agent development decision on this build. A planning agent reads each brief & splits the work between specialist agents. Each specialist then owns a single concern, such as the layout or the rules that tie objects together.

Their answers merge into one structured scene. One strict schema defines exactly what that scene may contain.

The model decides what goes where, & deterministic code turns that decision into geometry.

Our first decision was what the model would be allowed to produce. We ruled out free-form output early, because geometry code can’t negotiate with prose. Every agent answers inside the schema, & any other answer counts as a failed attempt.

Our second decision was where generation stops & code takes over. The model never emits a vertex or a mesh. It only describes the level, & an open-source mesh library builds real geometry from that description.

Every agent answers to the same contract, & that shared contract is what holds the whole chain together.

The model decides what goes where, & deterministic code turns that decision into geometry.

Architecture of the multi-agent AI level design pipeline, from one prompt through the schema checks to a level file THE PIPELINE LANGUAGE THE SCHEMA LINE PROMPT IN PLANNING AGENT LAYOUT AGENT OBJECT AGENT MERGED SCENE SCHEMA CHECK SPATIAL CHECK SENT BACK MESH BUILDER LEVEL FILE OPENS FILE GAME ENGINE HAND MODELING NOT IN THE LOOP Architecture of the multi-agent AI level design pipeline, stacked for small screens THE PIPELINE PROMPT IN PLANNING AGENT LAYOUT AGENT OBJECT AGENT MERGED SCENE SCHEMA CHECK SPATIAL CHECK SENT BACK MESH BUILDER LEVEL FILE OPENS FILE GAME ENGINE HAND MODELING NOT IN THE LOOP

Language stops at the schema line, & only numbers that pass validation ever become geometry.

The technology stack we used

Every layer in this stack exists to keep language apart from geometry. Our technology selection questions were the usual ones about what the model decides & what code guarantees. The list stays short on purpose, because fewer moving parts leave fewer places for a scene to go wrong.

Layer What we used Why What we ruled out
Language model A frontier model family Read spatial briefs more reliably than the others in our tests Two other families we benchmarked
Orchestration One agent per concern Small, focused prompts hold detail A single do-everything prompt
Scene schema One strict schema per scene Geometry code needs guarantees, not prose Free-form model output
Validation Schema & spatial checks Bad scenes bounce back before they become geometry Trusting the first answer
Mesh generation An open-source mesh library Deterministic geometry from validated numbers Asking the model for meshes
Output One standard 3D file Opens in the engines the team already uses Engine-locked exports

How does one prompt become a level?

One prompt becomes a level in six steps, in the exact order the pipeline runs them.

Six-step flow of one text prompt through agents, validation & mesh generation into a 3D level file LANGUAGE CODE 1 WRITE BRIEF 2 SPLIT BRIEF 3 GENERATE SCENE 4 VALIDATE SCENE SENT BACK 5 BUILD MESH 6 WRITE FILE Six steps from one written prompt to a level file, stacked for small screens ONE PROMPT, END TO END 1 · WRITE BRIEF 2 · SPLIT BRIEF 3 · GENERATE SCENE 4 · VALIDATE SCENE SENT BACK 5 · BUILD MESH 6 · WRITE FILE

A scene that fails a check goes straight back for repair before any geometry gets built.

  1. A designer writes the level in plain language, covering its spaces, paths, obstacles & overall mood.
  2. A planning agent reads the brief & splits it into separate jobs for the specialist agents downstream.
  3. Specialist agents turn each job into entries in one structured scene, with a set position & size for every object.
  4. A validation pass checks the scene against the schema & the spatial rules, then sends failures back for repair.
  5. The mesh stage reads the validated scene & builds real 3D geometry from it, one primitive for each entry.
  6. The pipeline writes one standard 3D file, & the team opens it in the engine to walk the level.

A person appears only at the two ends of that six-step chain.

What broke in the first builds?

Three problems broke the first builds, & they showed up in this order.

The first scenes read well but parsed badly. A field would go missing or a type would change, & some lists arrived as prose. The geometry code then crashed on input it had every right to trust.

The second failure was spatial rather than structural. Scenes began to pass validation & still made no sense, because objects overlapped or floated above the floor. Some walls sealed off the only path.

A scene could be valid JSON & still describe a level nobody could play.

The third problem was overload inside a single prompt. That one prompt handled the whole brief at once & dropped requirements as briefs grew longer. The model was stretched too thin for the job it had.

Research on language-guided 3D scene generation treats object positioning as a problem of its own[3]. The published fix gives a solver explicit spatial constraints instead of trusting a model’s raw coordinates. Weeks like these are when clients bring in specialist engineers instead of learning every lesson the slow way.

How we got past each one

We fixed each of the three problems separately, & every fix took several rounds to find.

Structure

We locked every agent to the schema & made every answer validate on arrival. A failed check goes back with the exact error, so the agent can repair its own output. Nothing malformed has reached the geometry code since that change landed.

Space

We added a spatial pass behind the schema pass. That pass checks for overlapping or floating objects & for spaces a player can’t reach. Each violation comes back with its coordinates, & a scene becomes geometry only after it clears both gates.

Split

We broke the single big prompt into the agent chain the pipeline runs today. Each agent holds one concern, so detail stopped falling out of long briefs.

Whitebox means the same thing as graybox, incidentally, & no one agrees on which name came first.

Surfacing this kind of work before anything gets promised is what an AI proof of concept is for.

A validation gate returning invalid scenes before mesh generation in the AI level pipeline SCHEMA CHECK SPATIAL CHECK TO MESH SENT BACK A validation gate returns invalid scenes before mesh generation, stacked for small screens THE VALIDATION GATE MODEL SCENE DOCUMENT SCHEMA CHECK SPATIAL CHECK TO MESH MESH SENT BACK FOR REPAIR

Every generated scene passes the schema gate & then the spatial gate, or it goes back for repair.

What changed after go-live?

Going live changed how every layout at the studio gets built, from the first idea to each revision.

What Before After
Getting to a walkable layout Hand placement in the editor One written prompt
Objects placed by hand Every object in the scene None
Revising a layout Rebuild the blockout Edit the prompt & run again
Trying variants One at a time, by hand One prompt each
What a designer does Builds the space Judges the space

We haven’t published a generation-time or hours-saved figure, because this build measured neither. We only print numbers we have actually measured on a build.

Day to day, AI level design changes where a designer’s attention goes. A layout idea now gets typed out instead of dragged into place. Every variant costs only one prompt, so exploring more ideas no longer means more building for the team.

The graybox itself stayed the same, & only the way it gets made has changed.

Planning spaces by hand before you build?

Tell us what your team still blocks out by hand, from game levels to warehouse racking. We’ll tell you where a written prompt could draft that layout instead.

What is running today

The AI level design pipeline is in use at the studio today, on the same agent chain the first version proved. A designer writes the brief & starts the run. The finished file then opens in the engine the team already uses. The schema is still the contract between language & geometry, & nothing becomes a mesh without passing it.

Every level the pipeline produces still starts as one prompt, just as it did on the first working run. Judging layouts is now the real design work, & typing the brief is the whole build step.

The running pipeline, with a written prompt passing through the agents, a structured scene & validation into an assembled graybox level 01 · PROMPT A ruined courtyard, walled on three sides, one gap leading east. Steps to a platform. Two pillars & a beacon. 02 · AGENTS PLANNING LAYOUT OBJECTS CONSTRAINTS 03 · SCENE slab x0 y0 18×12 wall x0 y12 18×1 wall x0 y0 1×12 steps x12 y4 3×3 pillar x3 y9 1×1 pillar x14 y9 1×1 04 · VALIDATE SCHEMA SPATIAL FAILED · SENT BACK FOR REPAIR 05 · LEVEL FILE WALKABLE ONE PROMPT IN ONE LEVEL FILE OUT The running pipeline, from prompt to walkable level, stacked for small screens THE PIPELINE, RUNNING 01 · PROMPT 02 · PLANNING AGENT LAYOUT, OBJECTS & CONSTRAINTS 03 · STRUCTURED SCENE 04 · SCHEMA CHECK SPATIAL CHECK SENT BACK 05 · LEVEL FILE WALKABLE

In the running pipeline, a failed scene goes back for repair before the level ever assembles.

What would we do differently?

Four changes would have saved us time on this build, & we now start every similar project with them.

Design the schema before the prompts

We wrote the prompts first & had to rework them once the schema landed. The schema turned out to be the product itself.

Put a check at every handoff

Our first validator sat at the very end, so one early mistake surfaced three stages late. Errors got cheap once checks ran at every handoff.

Walk every candidate before judging it

A generated level nobody can walk through is only a picture of a design. Early reviews leaned on top-down views, & clean layouts read wrong at player height. Now every candidate gets walked before anyone judges it.

Keep the model away from numbers

Letting the model fix coordinates is tempting when a scene is close. Every nudge invites drift, so repairs go back through the schema like everything else.

A generated level nobody can walk through is only a picture of a design.

Where else does prompt-to-3D fit?

Prompt-to-3D generation turns a written description into a walkable 3D layout, & it fits wherever space gets planned before it gets built.

Industry The equivalent problem What changes in the build
Construction Clients want walkable massing before drawings exist Real dimensions & code-driven constraint rules
Logistics warehouse rack layouts sketched before steel is ordered Racking sizes & aisle-width rules in the schema
Retail Store planners test fixture layouts for each site A fixture vocabulary & footfall constraints
Training simulation Safety teams need many scenario spaces at low cost Hazard placement rules & scenario variants
Virtual production Scenes blocked out before any set is built Camera-aware layouts & per-shot variants

Porting the pipeline takes a new object vocabulary in the schema plus the constraint rules of that domain, with a validation pass tuned to both.

A warehouse layout generated from a written prompt with the same 3D level pipeline SAME PIPELINE STRUCTURED PLAN NEW VOCABULARY A warehouse layout generated from a written prompt with the same pipeline, stacked for small screens WRITTEN BRIEF SAME PIPELINE STRUCTURED PLAN NEW VOCABULARY · RACKS & AISLES

The same prompt-to-3D pattern can plan a warehouse, with aisles & racking blocked out before anything is built.

Questions buyers usually ask

How it works

Can AI generate a 3D level from a text prompt?

Yes, as long as the model describes the level & code builds the geometry. A language model turns the prompt into a structured scene, & a mesh stage turns that scene into geometry. Asking a model for raw 3D geometry directly doesn’t hold up.

What is a graybox level, & why start there?

A graybox is a level built from plain untextured blocks to test layout & flow before any art exists. Teams start there because moving a gray block is cheap & moving finished art is expensive.

How is this different from procedural generation?

Classic procedural generation follows fixed rules or noise, so it can’t read a designer’s brief. A prompt-driven pipeline starts from written intent, & a designer who describes a level gets that level instead of a random one.

Time, cost & rollout

How long does a prompt-to-level pipeline take to build?

A working proof of concept on your own content usually takes a few weeks. The scene schema & the agent split each need a dedicated pass, & so do the validation rules. A production pipeline follows once generated levels start surviving review unedited.

How much does an AI level design build cost?

Cost depends on the variety of objects & the constraint rules, plus how much content you already have. Brainy Neurals scopes the build after a short conversation & a close look at your content, then quotes a fixed price. An AI readiness assessment tells you first whether your briefs & assets can drive one.

Does AI level generation replace level designers?

No, it replaces the placement labor & leaves the judgment about what plays well to your designers. They still decide what makes a space worth playing, & the pipeline gets them to that decision sooner.

What should a prompt build for you?

Levels, layouts, massing, racking & scenario spaces can all start as a prompt. If your team plans it before building it, tell us about the project & a written brief could probably draft it.







    Services behind this case study

    Brainy Neurals built this case study from the six services below.

    Generative AI development

    Text-to-3D pipelines built to return structured output instead of prose.

    AI agent development

    Multi-agent orchestration where specialist agents split a task & check each other’s work.

    RAG development

    Generation grounded in your own design documents & asset catalogs

    Document AI

    Messy briefs & design documents turned into structured data your tools can read.

    Hire AI developers

    Generative AI engineers who extend your team for the weeks when the schema fights back.

    AI in construction

    Layout & massing generation for construction, where space gets planned before it gets built.

    An AI proof of concept is the fastest way to test this on your own content. AI consulting helps you choose the model family & the line between language & geometry first. An AI readiness assessment tells you up front whether your briefs can drive a schema. The industries hub shows where else this generation pattern already runs.

    Similar case studies

    Brainy Neurals shipped each of these three builds into a real environment instead of a slide deck.

    Overhead Line Geometry Measurement

    Stereo cameras on a moving train measure wire geometry, with inference running on the train.

    AI Diet Assistant for Gastroenterology

    Clinical dietary guidance generated under review gates, live in a healthcare setting.

    Personalised AI Meal Planning for Chronic Care

    Structured meal plans built from messy personal health data, grounded & reviewable.

    Cite this case study

    Trivedi, Viral & Patel, Mitesh. AI Level Design That Builds Walkable Levels for Game Studios. Brainy Neurals, August 2026. https://brainyneurals.com/case-studies/ai-3d-level-design/

    Sources cited on this page

    1. Hendrikx M, Meijer S, Van der Velden J, Iosup A. Procedural Content Generation for Games: A Survey. ACM Transactions on Multimedia Computing, Communications, and Applications. 2013;9(1):1-22. DOI 10.1145/2422956.2422957.
    2. Summerville A, Snodgrass S, Guzdial M, Holmgård C, Hoover AK, Isaksen A, Nealen A, Togelius J. Procedural Content Generation via Machine Learning (PCGML). IEEE Transactions on Games. 2018;10(3):257-270. DOI 10.1109/TG.2018.2846639.
    3. Yang Y, Sun FY, Weihs L, VanderBilt E, Herrasti A, Han W, Wu J, Haber N, Krishna R, Liu L, Callison-Burch C, Yatskar M, Kembhavi A, Clark C. Holodeck: Language Guided Generation of 3D Embodied AI Environments. CVPR 2024. arXiv 2312.09067.