Home / Case studies / AI Business Simulation Platform for US Universities
Case study · EdTech & higher education · Multi-agent AI
AI Business Simulation Platform for US Universities
An EdTech company selling into US universities wanted an AI business simulation platform where students practise decisions & see what each one would cost. Brainy Neurals built a multi-tenant web platform that uses seven cooperating AI agents to run & score each business scenario in real time. Professors draft scenarios with AI help, & each university now runs in its own sealed tenant instead of the old one-university build. A third-party security audit found three cross-tenant leaks, & adversarial tests now prove on every deploy that they stay closed.
3 paths
Cross-tenant leaks closed
Per tenant
Daily & monthly AI budgets
Every deploy
Isolation tested again
Published October 2026
Key outcomes at a glance
The rebuilt platform delivered five outcomes that a university buyer can check for themselves.
Three leaks closed
A third-party security audit found three cross-tenant data paths. Each one is closed & covered by a test that failed before the fix & passes after it.
A full move to Python
More than 42 API routes & a 26-task agent pipeline moved from TypeScript to Python. Six milestones checked each piece against the original.
Seven agents per turn
Seven specialized AI agents score each student decision against a teaching rubric in real time.
Per-tenant AI budgets
Per-tenant AI budgets with daily & monthly caps are enforced live with a real-time cost dashboard, replacing zero prior spend controls.
An audit trail for FERPA
Every read, write, delete & admin action on student data is logged across all tenants. Those logs are kept for at least two years.
Why do business schools need simulations?
Business schools need simulations because students rarely see what their decisions would do to a real company. A student reads a case, argues a position in class & never learns how that choice would hit stock or cash.
Brainy Neurals built an AI business simulation platform for an EdTech company selling into US higher education. In each branching scenario, decisions carry consequences & an AI counterpart pushes back on the student. A rubric then scores the reasoning behind every choice.
The first problem was authoring, because one realistic simulation took a very long time to build by hand. As a result, catalogs grew slowly & went stale.
Visibility was the second problem, since professors saw final answers but never the reasoning behind them. They couldn’t tell who applied Porter’s Five Forces & who only name-dropped it.
Cost was the third problem, because nobody could predict what a full cohort would spend on language models. Nothing in the old platform metered that spend.
The engagement asked for all of this inside one production platform. Our generative AI development scope covered the agents & the infrastructure they run on.
What stopped the first platform from scaling?
The first version of the AI business simulation platform proved the teaching idea but could not scale beyond one school. It ran as one TypeScript & Node.js application with a single database & no tenant boundary anywhere in the code.
Onboarding a second university would have meant copying the whole stack or sharing tables with nothing between them.
Four gaps in tenancy, audit, spend control & toolchain kept the original monolith to one university.
Four gaps made the old platform impossible to ship to more schools. There was no multi-tenant isolation, so one missed filter could expose a school’s student records to another school.
Student data also had no audit trail, which FERPA deployment reviews in US higher education ask for. AI spend had no controls either, so every new user pushed an open-ended model bill higher.
Finally, the Node.js runtime sat awkwardly under an AI workload that the Python ecosystem serves better. Evaluation tools, data libraries, PDF generation & the agent roadmap all pointed toward Python. We chose a structured replatform over a quick patch.
How does the AI business simulation platform work?
Every simulation turn runs through seven specialized agents, each with one job & a defined contract. That pipeline is the core of our AI agent development work on this engagement.
Seven specialized agents share typed contracts under one FastAPI layer, & every output lands in PostgreSQL with full lineage.
The Director sequences each scenario, choosing the next decision point & when to end. Its brief goes to the Narrator, which turns it into what the student sees, like a supplier email or a board memo. Scoring falls to the Evaluator, which checks each decision against a teaching competency rubric.
A Domain Expert grounds every reply in the subject, so a supply-chain scenario behaves like real supply chain work. Reasoning depth is measured by the Depth Evaluator across three strictness modes. When a student asks why something happened, the Causal Explainer traces the result back to the decision behind it.
On the professor’s side, the Authoring Assistant turns a plain description of a business situation into a structured simulation with several decisions. Professors edit that draft instead of starting from a blank page.
Two design rules hold the pipeline together. Agents talk through typed, versioned contracts, so one prompt change can’t quietly break downstream scoring.
Every agent output also lands in PostgreSQL with full lineage, which makes class-wide analytics possible. The platform spots strategic frameworks such as Theory of Constraints & Porter’s Five Forces in scenarios & in live student reasoning. It then sums up those patterns for each cohort.
What happens when a student makes a decision?
When a student makes a decision, such as cutting the marketing budget, the Director receives it with the full scenario state. It decides what the world does next & hands the Narrator a brief.
While the Narrator writes the consequence, the Evaluator & Depth Evaluator score the decision in parallel. One track checks rubric skills & the other checks reasoning depth, with strictness set per course.
Meanwhile the Causal Explainer prepares the reason behind the outcome. Live business numbers such as revenue & cash then update on screen.
Scoring runs in parallel lanes on every turn, so students see a reacting world while professors get the data.
None of the scoring interrupts the conversation with the student. Students see a world that reacts, while professors see the measurements. After the final decision, the platform writes a structured debrief & a full PDF transcript of the session.
Concurrency matters just as much in a real classroom. Dozens of students hit the same pipeline within the same 10 minutes, so turns queue through Redis & responses cache where results repeat.
An OpenRouter layer falls back between model providers & rotates across several keys, so one provider outage can’t end a lecture. Its routes include Anthropic & Gemini models, & OpenAI models sit on the same routing list.
We first hardened that queue & fallback pattern on our multi-agent browser automation case study.
Six milestones from TypeScript to Python
The TypeScript to Python migration ran in six milestones, because rewrites fail when they try to do everything at once. Each milestone ended in a runnable system with parity checks against the TypeScript original.
The order was the database layer, then authentication & role-based access, then the 42+ API routes. After that came the 26-task agent pipeline, followed by analytics & exports. The multi-tenant re-architecture came last, as the goal the whole exercise existed for.
Each of the six milestones closed only when the Python routes returned the same shapes as the TypeScript originals.
Parity was the discipline behind every milestone. Each migrated route had to return the same shapes for the same inputs before its milestone closed. A milestone closed when the Python routes answered like the TypeScript routes, & not before.
SQLAlchemy 2.0 with Alembic replaced the ad hoc schema of the original. Today the whole platform rebuilds from an empty PostgreSQL database through migrations alone.
Cloud Build deploys it with no manual or undocumented steps. Secret Manager holds every credential, so nothing ships in an environment file. One team carried the six milestones end to end.
A milestone closed when the Python routes answered like the TypeScript routes, & not before.
How do you prove tenant data stays separate?
You prove tenant data stays separate by attacking it, because multi-tenancy fails quietly. Nothing crashes when a query misses its tenant filter. The wrong rows simply come back, & on an education platform those rows are student records.
OWASP ranks broken access control as the top web application risk, & multi-tenant SaaS is exactly where that risk bites. So we treated isolation as a claim that needs outside proof & sent the rebuilt platform to a third-party security audit.
The audit surfaced three cross-tenant data access paths. Each fix started with an adversarial test written from the attacker’s side. The test had to fail on the old build before the fix shipped, & the same test had to pass after it.
Those adversarial tests now live in CI for good. Every deploy proves again that one tenant can’t read another tenant’s data, so tenant isolation is a regression suite rather than a launch-day memory. An isolation claim you haven’t attacked is only a hope.
The three cross-tenant paths a third-party audit surfaced are now adversarial tests that run on every deploy.
The same design carries the FERPA compliance load. Every read, write, delete & admin action on student data writes an audit event with the actor, tenant, time & object.
Deletes & admin actions are included, which is where audit trails often go thin. Events are kept for at least two years to meet FERPA review in US higher education.
Role-based access covers Student, Professor, Admin & Super Admin roles. A Switch View mode lets admins preview exactly what another role sees without holding that role’s permissions. We took this audit-first habit from AI tender document analysis, an earlier build where provenance mattered as much as answers.
An isolation claim you haven’t attacked is only a hope.
How are AI costs capped per tenant?
AI costs are capped per tenant inside the platform, because every simulation turn spends money. Attribution came first, so every language model call carries tenant, user, agent & task details into a usage ledger.
Each call is priced against live provider price tables behind OpenRouter. The dashboard reads that ledger in real time & splits spend by tenant & by agent. A university admin & the platform team look at the same number.
Every tenant carries its own daily & monthly caps, & the dashboard reads the same ledger that enforcement uses. All figures shown in this dashboard are illustrative.
Enforcement is the step that turns metering into real control. Each tenant has configurable daily & monthly spend caps. Approaching a cap raises alerts, & hitting one degrades service deliberately instead of failing it.
A cap becomes a controlled slowdown rather than an outage. Nobody finds the overrun on an invoice weeks later, because the platform refuses to let it get there.
Provider routing serves the same goal of predictable cost. OpenRouter adds automatic fallback & multi-key rotation, which keeps uptime & unit costs inside chosen bounds. Before this system, the platform had zero spend controls of any kind.
What runs in production today?
Today the platform runs in production as a true multi-tenant SaaS on Google Cloud Run. PostgreSQL sits behind the Cloud SQL Auth Proxy, & Firebase Authentication handles Google Sign-In & role propagation.
Each university is a fully isolated tenant, so onboarding another institution is configuration rather than an engineering project.
Professors write scenarios with AI help & publish them to their own catalog. Students run branching simulations that are scored in real time. Professor analytics now show what used to be invisible, such as cohort decision patterns & module health. They also track depth trends & per-student reasoning signals, down to the frameworks a class really used.
The whole product ships in English & Spanish, including AI-generated content. Session transcripts export to PDF, & Redis-backed queueing holds classroom concurrency under live load.
The replatform swapped untested boundaries & unmetered spend for proven isolation & full audit coverage, with budgets enforced per tenant.
The business result is a single-university pilot that became an institution-ready platform. The isolation proof & audit trail that university procurement reviews ask about now live in the architecture instead of a slide.
Our real-time voice AI English tutor build follows the same pattern, because education products live or die under live classroom load.
Serving many schools from one AI platform?
Tell us where your platform stands today, or check how ready your team is for AI first. Both routes start from your stack & your compliance needs.
Before & after, dimension by dimension
The table compares the old monolith & the new platform across seven dimensions.
| Dimension | Single-tenant monolith (before) | Multi-tenant platform (after) |
|---|---|---|
| Tenancy & onboarding | One university, hard-coded, so a second customer meant a stack copy | Isolated tenants, with onboarding done through configuration |
| Data-isolation proof | None, with shared code paths & untested boundaries | Three audit findings closed, with adversarial tests on every deploy |
| Audit coverage | No access log on student data | Every read, write, delete & admin action logged, kept two years |
| AI spend control | None, so adoption grew the bill without limit | Per-tenant daily & monthly caps with a live dashboard |
| Scenario authoring | Fully manual instructional design | AI-assisted, with a description in & a structured draft out |
| Instructor visibility | Final grades only | Cohort patterns, framework detection & depth trends |
| Rebuild & deploys | Manual, undocumented steps | Alembic & Cloud Build, with clean rebuilds & CI/CD deploys |
The ROI math
The ROI math runs on three lines, & we show the working instead of the adjectives.
Authoring time
AI-assisted authoring turns scenario creation from a manual design task into an editable draft that a professor refines. That saving compounds across every scenario in every course.
AI spend
Caps turn language model cost from an open-ended risk into a budget line, so course pricing can be worked out in advance.
Market access
Single-tenant architecture priced every new university as an engineering project. Multi-tenant onboarding is configuration, & the third-party isolation proof shortens the security review that often slows higher-education sales.
What we’d tell you before you build
We would share five lessons & one honest cost with any team planning a platform like this one.
Parity gates beat big-bang rewrites
Six milestones with hard parity checks kept the migration close to a demo-ready state, & rollback stayed cheap.
Buy the audit before customers force it
The third-party security review found three leaks we then proved closed. Finding them inside a university’s procurement review would have cost the deal.
Make isolation a regression suite
Adversarial tests that fail before a fix & pass after are the only isolation proof that survives the next deploy.
Meter before you optimize
Per-tenant attribution had to exist before caps meant anything, since you can’t enforce a budget you can’t see.
Audit logs earn their keep on deletes
Read & write logging is common, but delete & admin-action logging is where FERPA reviews probe. The original platform had none of it.
One honest cost
Typed agent contracts slowed the first milestone while the team argued over schemas. Every later milestone repaid it, because pipeline changes stopped breaking downstream scoring.
A rewrite you can’t demo mid-flight is a rewrite you can’t steer, so every one of the six milestones stayed shippable.
A rewrite you can’t demo mid-flight is a rewrite you can’t steer, so every one of the six milestones stayed shippable.
Could this run your platform?
Your platform can use the same pattern if it runs AI interactions for many customer organizations, such as training or tutoring products. Assessment, coaching & any regulated multi-tenant SaaS fit as well.
The same four moves apply in each case. They are typed agent contracts & tenant isolation proven by adversarial tests. Per-tenant spend enforcement & audit logging built for the review that will read it complete the set.
The migration playbook works even without AI, because any single-tenant platform facing its second enterprise customer faces the same six-milestone shape.
Scoping your version takes three answers from you.
- Your current stack & tenancy model
- The compliance regime your buyers cite, whether FERPA, GDPR or SOC 2
- Where AI spend goes unmetered today
With those answers we can usually propose a realistic engagement structure & a budget range.
Where else does this fit?
A multi-tenant AI simulation platform runs scored, branching decision scenarios for many organizations at once. Shared infrastructure carries all of them, & hard tenant boundaries keep each organization’s data its own. Spend & audit trails stay separate the same way.
Business simulation was the first place we shipped this. The same design moves anywhere decisions carry consequences & a reviewer needs to see the reasoning behind the answer. Five sectors ask for the same engine with the domain swapped out. The generative AI development scope stays the same across all of them.
| Sector | The equivalent problem | What changes in the build |
|---|---|---|
| Corporate learning & development | Compliance & leadership training tests recall rather than decisions under pressure | Swap the rubric for the competency model HR already runs |
| Healthcare clinical training | Nurses & residents rehearse rare cases too dangerous to practice live | Clinical review gates on content, with stricter audit under health-record rules |
| Financial services training | Advisors learn suitability & disclosure from slides, then face real clients | A scenario library tied to the regime, with scoring against conduct rules |
| Aviation & defense training | Crews drill emergency procedures where a wrong call cascades fast | Tighter scenario branching, with evaluation mapped to the certifying body’s framework |
| Government & public-sector training | Caseworkers make judgment calls the agency must later defend | Longer audit retention, with data residency inside the agency’s own boundary |
Porting needs a new scenario library & a rubric the sector trusts, set inside the buyer’s compliance regime. The agents, the isolation proof, the spend caps & the audit spine carry over unchanged.
Where this platform goes next
Next, this platform becomes our reference pattern for regulated multi-tenant AI products. The agents change with each domain, while the isolation proof & spend enforcement stay the same.
The audit spine carries over to every new domain as well. If your platform is one tenant away from needing all of this, the architecture in this write-up is where we would start.
You can read more about how we work on the About Brainy Neurals page.
Questions buyers actually ask
Buyers tend to ask these questions when they scope a multi-tenant AI platform.
How much does it cost to build an AI business simulation platform?
A production platform of this shape, with a multi-agent engine, multi-tenancy & compliance logging, is usually a multi-month build. Agent count & evaluation depth move the number, along with the tenancy model & the compliance regime. How much of an existing platform migrates or retires matters too. Sharing your constraints through the form gets you a range faster than a spreadsheet of assumptions.
Can AI reliably score student reasoning?
Yes, reliably enough for formative assessment, as long as the scoring is engineered. This platform scores against an explicit competency rubric with three strictness modes & tracks reasoning depth apart from correctness. Every evaluation is logged for professor review, & framework detection flags what students applied rather than merely mentioned. Professors keep final grading authority while the AI supplies evidence at a scale humans can’t match.
What makes an AI education platform FERPA-compliant?
FERPA compliance comes from architecture plus contract. On the architecture side you need access controls on student records & audit logging of reads, writes, deletes & admin actions. You also need a retention & deletion policy, & student data must stay out of model training. The contract side is the school-official designation with explicit data-use limits, as the US Department of Education’s student privacy guidance explains. EU deployments face a parallel review under GDPR & the EU AI Act, while Australian ones face the Privacy Act.
How do you test multi-tenant data isolation?
You test it adversarially, with tests written from the attacker’s side. A good suite has tenant A request tenant B’s objects by ID & cross scope through joins. It also probes the admin screens for scope leaks. Each test must fail before a fix ships & pass after it, & the suite runs on every deploy. A third-party audit found three such paths here, & the adversarial suite is why they stayed closed.
Shared schema or schema per tenant, which fits?
Shared schema with enforced tenant scoping fits many SaaS products until compliance pressure gets serious. It gives you one migration path & one connection pool at the lowest cost. It also leaves you one missed filter away from a leak, which is why adversarial testing matters so much there. Schema per tenant or database per tenant buys stronger boundaries at a real operating cost in migrations & backups. Our single-tenant to multi-tenant migration playbook walks through that decision.
How do you keep LLM costs predictable as adoption grows?
Start with attribution & add enforcement once spend is visible. Tag every model call with its tenant & its user at request time. Record the task too, then price each call against live provider tables in a real-time ledger. Next, give every tenant daily & monthly caps with alerts & a deliberate slowdown path when a cap hits. Multi-provider routing keeps unit costs inside chosen bounds, & costs stay predictable because the platform refuses overruns instead of reporting them later.
Why move from Node.js to Python for AI workloads?
The AI ecosystem centers on Python, from evaluation tools & data libraries to agent frameworks & the hiring pool around them. FastAPI with SQLAlchemy 2.0 gave this platform typed contracts & async speed close to the Node original, while opening the toolchain the roadmap needed. The migration cost is real, & six parity-gated milestones kept it manageable.
Tell us what you're building
Share your current stack & the compliance rules your buyers cite. The person who would architect your platform reads every message & replies.
Services behind this case study
One Brainy Neurals service line delivered this build from the agents down to the infrastructure.
Generative AI development
Multi-agent systems & evaluation layers built to run in production for many tenants.
Similar case studies
Similar case studies show the same patterns under different workloads.
Real-time voice AI English tutor
Sub-second conversational tutoring built to survive live classroom load.
Multi-agent browser automation
Where we first hardened the queue & fallback pattern behind this engine.
AI tender document analysis
The audit-first, provenance-heavy habit we carried into isolation proof.
You can browse every build in Brainy Neurals case studies.
Cite this case study
Cite this case study with the reference below, which names the page & its permanent address.
Brainy Neurals (2026). AI Business Simulation Platform for US Universities. Brainy Neurals case study. https://www.brainyneurals.com/case-studies/ai-business-simulation-platform










