AI Business Simulation Platform for US Universities

Home / Case studies / AI Business Simulation Platform for US Universities

Case study · EdTech & higher education · Multi-agent AI

AI Business Simulation Platform for US Universities

An EdTech company selling into US universities wanted an AI business simulation platform where students practise decisions & see what each one would cost. Brainy Neurals built a multi-tenant web platform that uses seven cooperating AI agents to run & score each business scenario in real time. Professors draft scenarios with AI help, & each university now runs in its own sealed tenant instead of the old one-university build. A third-party security audit found three cross-tenant leaks, & adversarial tests now prove on every deploy that they stay closed.

  • #AIBusinessSimulation
  • #MultiAgentAI
  • #HigherEducation
  • #EdTech
  • #TenantIsolation
  • #FERPACompliance

3 paths

Cross-tenant leaks closed

Per tenant

Daily & monthly AI budgets

Every deploy

Isolation tested again

Ronak Patel
Preksha Rana

Published October 2026

Key outcomes at a glance

The rebuilt platform delivered five outcomes that a university buyer can check for themselves.

Three leaks closed

A third-party security audit found three cross-tenant data paths. Each one is closed & covered by a test that failed before the fix & passes after it.

A full move to Python

More than 42 API routes & a 26-task agent pipeline moved from TypeScript to Python. Six milestones checked each piece against the original.

Seven agents per turn

Seven specialized AI agents score each student decision against a teaching rubric in real time.

Per-tenant AI budgets

Per-tenant AI budgets with daily & monthly caps are enforced live with a real-time cost dashboard, replacing zero prior spend controls.

An audit trail for FERPA

Every read, write, delete & admin action on student data is logged across all tenants. Those logs are kept for at least two years.

Why do business schools need simulations?

Business schools need simulations because students rarely see what their decisions would do to a real company. A student reads a case, argues a position in class & never learns how that choice would hit stock or cash.

Brainy Neurals built an AI business simulation platform for an EdTech company selling into US higher education. In each branching scenario, decisions carry consequences & an AI counterpart pushes back on the student. A rubric then scores the reasoning behind every choice.

The first problem was authoring, because one realistic simulation took a very long time to build by hand. As a result, catalogs grew slowly & went stale.

Visibility was the second problem, since professors saw final answers but never the reasoning behind them. They couldn’t tell who applied Porter’s Five Forces & who only name-dropped it.

Cost was the third problem, because nobody could predict what a full cohort would spend on language models. Nothing in the old platform metered that spend.

The engagement asked for all of this inside one production platform. Our generative AI development scope covered the agents & the infrastructure they run on.

What stopped the first platform from scaling?

The first version of the AI business simulation platform proved the teaching idea but could not scale beyond one school. It ran as one TypeScript & Node.js application with a single database & no tenant boundary anywhere in the code.

Onboarding a second university would have meant copying the whole stack or sharing tables with nothing between them.

Four gaps in the single-tenant monolith MONOLITH · V1 TypeScript / Node.js one database · one university shared tables no tenant column No multi-tenant isolation a missed filter exposes another school’s records No audit trail on student data exactly what FERPA deployment reviews ask for No AI spend controls adoption growth means an unbounded API bill AI toolchain mismatch the Python ecosystem serves the roadmap better Second university onboarding blocked without a replatform BLOCKED Four gaps in the single-tenant monolith MONOLITH · V1 TypeScript / Node.js one database · one university No multi-tenant isolation a missed filter leaks records No audit trail FERPA reviews ask for one No AI spend controls an open-ended API bill AI toolchain mismatch Python serves the roadmap Second university onboarding blocked

Four gaps in tenancy, audit, spend control & toolchain kept the original monolith to one university.

Four gaps made the old platform impossible to ship to more schools. There was no multi-tenant isolation, so one missed filter could expose a school’s student records to another school.

Student data also had no audit trail, which FERPA deployment reviews in US higher education ask for. AI spend had no controls either, so every new user pushed an open-ended model bill higher.

Finally, the Node.js runtime sat awkwardly under an AI workload that the Python ecosystem serves better. Evaluation tools, data libraries, PDF generation & the agent roadmap all pointed toward Python. We chose a structured replatform over a quick patch.

How does the AI business simulation platform work?

Every simulation turn runs through seven specialized agents, each with one job & a defined contract. That pipeline is the core of our AI agent development work on this engagement.

Seven-agent architecture of the platform Student Professor FIREBASE AUTH Google Sign-In · JWT roles ORCHESTRATION FastAPI · Python 3.12 · typed, versioned agent contracts Director sequences Narrator the world Evaluator rubric Domain Expert grounds Depth Evaluator 3 modes Causal Explainer the why Authoring Assistant professors POSTGRESQL lineage + rubric scores REDIS queue + response cache OpenRouter Anthropic Gemini OpenAI automatic fallback + multi-key rotation RUNS ON · Google Cloud Run · Cloud Build CI/CD · Secret Manager Seven-agent architecture of the platform Student · Professor Firebase Auth sign-in FastAPI · Python 3.12 typed agent contracts Seven agents per turn Director · Narrator Evaluator · Domain Expert Depth Evaluator Causal Explainer Authoring Assistant PostgreSQL · Redis lineage · queue · cache OpenRouter fallback · key rotation RUNS ON GOOGLE CLOUD RUN

Seven specialized agents share typed contracts under one FastAPI layer, & every output lands in PostgreSQL with full lineage.

The Director sequences each scenario, choosing the next decision point & when to end. Its brief goes to the Narrator, which turns it into what the student sees, like a supplier email or a board memo. Scoring falls to the Evaluator, which checks each decision against a teaching competency rubric.

A Domain Expert grounds every reply in the subject, so a supply-chain scenario behaves like real supply chain work. Reasoning depth is measured by the Depth Evaluator across three strictness modes. When a student asks why something happened, the Causal Explainer traces the result back to the decision behind it.

On the professor’s side, the Authoring Assistant turns a plain description of a business situation into a structured simulation with several decisions. Professors edit that draft instead of starting from a blank page.

Two design rules hold the pipeline together. Agents talk through typed, versioned contracts, so one prompt change can’t quietly break downstream scoring.

Every agent output also lands in PostgreSQL with full lineage, which makes class-wide analytics possible. The platform spots strategic frameworks such as Theory of Constraints & Porter’s Five Forces in scenarios & in live student reasoning. It then sums up those patterns for each cohort.

What happens when a student makes a decision?

When a student makes a decision, such as cutting the marketing budget, the Director receives it with the full scenario state. It decides what the world does next & hands the Narrator a brief.

While the Narrator writes the consequence, the Evaluator & Depth Evaluator score the decision in parallel. One track checks rubric skills & the other checks reasoning depth, with strictness set per course.

Meanwhile the Causal Explainer prepares the reason behind the outcome. Live business numbers such as revenue & cash then update on screen.

One simulation turn from decision to KPI update Student decision Director next world state Narrator the consequence shown Evaluator rubric competencies Depth Evaluator 3 strictness modes Causal Explainer the why KPIs update revenue · cash service level Debrief + PDF ReportLab transcript loop · next decision One simulation turn from decision to KPI update Student decision Director next world state SCORED IN PARALLEL Narrator the consequence shown Evaluator rubric skills Depth Evaluator three strictness modes Causal Explainer the why KPIs update revenue · cash · service Next decision or debrief PDF transcript at the end

Scoring runs in parallel lanes on every turn, so students see a reacting world while professors get the data.

None of the scoring interrupts the conversation with the student. Students see a world that reacts, while professors see the measurements. After the final decision, the platform writes a structured debrief & a full PDF transcript of the session.

Concurrency matters just as much in a real classroom. Dozens of students hit the same pipeline within the same 10 minutes, so turns queue through Redis & responses cache where results repeat.

An OpenRouter layer falls back between model providers & rotates across several keys, so one provider outage can’t end a lecture. Its routes include Anthropic & Gemini models, & OpenAI models sit on the same routing list.

We first hardened that queue & fallback pattern on our multi-agent browser automation case study.

Six milestones from TypeScript to Python

The TypeScript to Python migration ran in six milestones, because rewrites fail when they try to do everything at once. Each milestone ended in a runnable system with parity checks against the TypeScript original.

The order was the database layer, then authentication & role-based access, then the 42+ API routes. After that came the 26-task agent pipeline, followed by analytics & exports. The multi-tenant re-architecture came last, as the goal the whole exercise existed for.

Six-milestone migration from TypeScript to Python TYPESCRIPT / NODE.JS PYTHON · FASTAPI M1 Database layer parity passed M2 Auth + RBAC parity passed M3 42+ API routes parity passed M4 26-task agent pipeline parity passed M5 Analytics + exports parity passed M6 Multi-tenant re-architecture parity passed ✓ Parity checks passed at every gate, with Python routes matching the TypeScript originals Six-milestone migration from TypeScript to Python M1 · Database layer parity passed M2 · Auth + RBAC parity passed M3 · 42+ API routes parity passed M4 · 26-task agent pipeline parity passed M5 · Analytics + exports parity passed M6 · Multi-tenant rebuild parity passed

Each of the six milestones closed only when the Python routes returned the same shapes as the TypeScript originals.

Parity was the discipline behind every milestone. Each migrated route had to return the same shapes for the same inputs before its milestone closed. A milestone closed when the Python routes answered like the TypeScript routes, & not before.

SQLAlchemy 2.0 with Alembic replaced the ad hoc schema of the original. Today the whole platform rebuilds from an empty PostgreSQL database through migrations alone.

Cloud Build deploys it with no manual or undocumented steps. Secret Manager holds every credential, so nothing ships in an environment file. One team carried the six milestones end to end.

A milestone closed when the Python routes answered like the TypeScript routes, & not before.

How do you prove tenant data stays separate?

You prove tenant data stays separate by attacking it, because multi-tenancy fails quietly. Nothing crashes when a query misses its tenant filter. The wrong rows simply come back, & on an education platform those rows are student records.

OWASP ranks broken access control as the top web application risk, & multi-tenant SaaS is exactly where that risk bites. So we treated isolation as a claim that needs outside proof & sent the rebuilt platform to a third-party security audit.

The audit surfaced three cross-tenant data access paths. Each fix started with an adversarial test written from the attacker’s side. The test had to fail on the old build before the fix shipped, & the same test had to pass after it.

Those adversarial tests now live in CI for good. Every deploy proves again that one tenant can’t read another tenant’s data, so tenant isolation is a regression suite rather than a launch-day memory. An isolation claim you haven’t attacked is only a hope.

Tenant isolation proven by adversarial tests TENANT · UNIVERSITY A Users (Student · Prof) Tenant-scoped queries Tenant A data TENANT · UNIVERSITY B Users (Student · Prof) Tenant-scoped queries Tenant B data tenant boundary ADVERSARIAL TEST A requests B’s object FAIL pre-fix PASS post-fix Audit findings fixed, then re-tested in CI on every deploy isolation = regression suite Tenant isolation proven by adversarial tests UNIVERSITY A Tenant A data tenant-scoped queries TENANT BOUNDARY Adversarial test A requests B’s object FAIL on the pre-fix build PASS on the post-fix build UNIVERSITY B Tenant B data tenant-scoped queries Tests run in CI on every deploy

The three cross-tenant paths a third-party audit surfaced are now adversarial tests that run on every deploy.

The same design carries the FERPA compliance load. Every read, write, delete & admin action on student data writes an audit event with the actor, tenant, time & object.

Deletes & admin actions are included, which is where audit trails often go thin. Events are kept for at least two years to meet FERPA review in US higher education.

Role-based access covers Student, Professor, Admin & Super Admin roles. A Switch View mode lets admins preview exactly what another role sees without holding that role’s permissions. We took this audit-first habit from AI tender document analysis, an earlier build where provenance mattered as much as answers.

An isolation claim you haven’t attacked is only a hope.

How are AI costs capped per tenant?

AI costs are capped per tenant inside the platform, because every simulation turn spends money. Attribution came first, so every language model call carries tenant, user, agent & task details into a usage ledger.

Each call is priced against live provider price tables behind OpenRouter. The dashboard reads that ledger in real time & splits spend by tenant & by agent. A university admin & the platform team look at the same number.

Per-tenant AI spend dashboard with illustrative figures Per-tenant AI spend · live ledger ILLUSTRATIVE FIGURES reads the usage ledger daily cap monthly Northgate University 54% Lakeside Business School 92% approaching daily cap Meridian College 27% Cap hit means a deliberate slowdown alerts on approach · in-flight turns preserved PROVIDER SPLIT Anthropic 45% Gemini 30% OpenAI 25% Per-tenant AI spend dashboard with illustrative figures ILLUSTRATIVE FIGURES Northgate University 54% of cap Lakeside Business School 92% · near daily cap Meridian College 27% of cap Provider split Anthropic 45% Gemini 30% · OpenAI 25% Cap hit service slows on purpose

Every tenant carries its own daily & monthly caps, & the dashboard reads the same ledger that enforcement uses. All figures shown in this dashboard are illustrative.

Enforcement is the step that turns metering into real control. Each tenant has configurable daily & monthly spend caps. Approaching a cap raises alerts, & hitting one degrades service deliberately instead of failing it.

A cap becomes a controlled slowdown rather than an outage. Nobody finds the overrun on an invoice weeks later, because the platform refuses to let it get there.

Provider routing serves the same goal of predictable cost. OpenRouter adds automatic fallback & multi-key rotation, which keeps uptime & unit costs inside chosen bounds. Before this system, the platform had zero spend controls of any kind.

What runs in production today?

Today the platform runs in production as a true multi-tenant SaaS on Google Cloud Run. PostgreSQL sits behind the Cloud SQL Auth Proxy, & Firebase Authentication handles Google Sign-In & role propagation.

Each university is a fully isolated tenant, so onboarding another institution is configuration rather than an engineering project.

Professors write scenarios with AI help & publish them to their own catalog. Students run branching simulations that are scored in real time. Professor analytics now show what used to be invisible, such as cohort decision patterns & module health. They also track depth trends & per-student reasoning signals, down to the frameworks a class really used.

The whole product ships in English & Spanish, including AI-generated content. Session transcripts export to PDF, & Redis-backed queueing holds classroom concurrency under live load.

Before & after capabilities of the platform BEFORE · MONOLITH Single tenant, hard-coded No isolation proof No audit log on student data Unmetered AI spend Manual scenario authoring AFTER · MULTI-TENANT PLATFORM Isolated tenants Adversarial tests in CI Every access logged · 2-year floor Per-tenant daily / monthly caps AI-assisted authoring Before & after capabilities of the platform Before · monolith single tenant, hard-coded no isolation proof no audit log unmetered AI spend manual authoring After · multi-tenant isolated tenants adversarial tests in CI every access logged per-tenant caps AI-assisted authoring

The replatform swapped untested boundaries & unmetered spend for proven isolation & full audit coverage, with budgets enforced per tenant.

The business result is a single-university pilot that became an institution-ready platform. The isolation proof & audit trail that university procurement reviews ask about now live in the architecture instead of a slide.

Our real-time voice AI English tutor build follows the same pattern, because education products live or die under live classroom load.

Serving many schools from one AI platform?

Tell us where your platform stands today, or check how ready your team is for AI first. Both routes start from your stack & your compliance needs.

Before & after, dimension by dimension

The table compares the old monolith & the new platform across seven dimensions.

Dimension Single-tenant monolith (before) Multi-tenant platform (after)
Tenancy & onboarding One university, hard-coded, so a second customer meant a stack copy Isolated tenants, with onboarding done through configuration
Data-isolation proof None, with shared code paths & untested boundaries Three audit findings closed, with adversarial tests on every deploy
Audit coverage No access log on student data Every read, write, delete & admin action logged, kept two years
AI spend control None, so adoption grew the bill without limit Per-tenant daily & monthly caps with a live dashboard
Scenario authoring Fully manual instructional design AI-assisted, with a description in & a structured draft out
Instructor visibility Final grades only Cohort patterns, framework detection & depth trends
Rebuild & deploys Manual, undocumented steps Alembic & Cloud Build, with clean rebuilds & CI/CD deploys

The ROI math

The ROI math runs on three lines, & we show the working instead of the adjectives.

Authoring time

AI-assisted authoring turns scenario creation from a manual design task into an editable draft that a professor refines. That saving compounds across every scenario in every course.

AI spend

Caps turn language model cost from an open-ended risk into a budget line, so course pricing can be worked out in advance.

Market access

Single-tenant architecture priced every new university as an engineering project. Multi-tenant onboarding is configuration, & the third-party isolation proof shortens the security review that often slows higher-education sales.

What we’d tell you before you build

We would share five lessons & one honest cost with any team planning a platform like this one.

Parity gates beat big-bang rewrites

Six milestones with hard parity checks kept the migration close to a demo-ready state, & rollback stayed cheap.

Buy the audit before customers force it

The third-party security review found three leaks we then proved closed. Finding them inside a university’s procurement review would have cost the deal.

Make isolation a regression suite

Adversarial tests that fail before a fix & pass after are the only isolation proof that survives the next deploy.

Meter before you optimize

Per-tenant attribution had to exist before caps meant anything, since you can’t enforce a budget you can’t see.

Audit logs earn their keep on deletes

Read & write logging is common, but delete & admin-action logging is where FERPA reviews probe. The original platform had none of it.

One honest cost

Typed agent contracts slowed the first milestone while the team argued over schemas. Every later milestone repaid it, because pipeline changes stopped breaking downstream scoring.

A rewrite you can’t demo mid-flight is a rewrite you can’t steer, so every one of the six milestones stayed shippable.

A rewrite you can’t demo mid-flight is a rewrite you can’t steer, so every one of the six milestones stayed shippable.

Could this run your platform?

Your platform can use the same pattern if it runs AI interactions for many customer organizations, such as training or tutoring products. Assessment, coaching & any regulated multi-tenant SaaS fit as well.

The same four moves apply in each case. They are typed agent contracts & tenant isolation proven by adversarial tests. Per-tenant spend enforcement & audit logging built for the review that will read it complete the set.

The migration playbook works even without AI, because any single-tenant platform facing its second enterprise customer faces the same six-milestone shape.

Scoping your version takes three answers from you.

  • Your current stack & tenancy model
  • The compliance regime your buyers cite, whether FERPA, GDPR or SOC 2
  • Where AI spend goes unmetered today

With those answers we can usually propose a realistic engagement structure & a budget range.

Where else does this fit?

A multi-tenant AI simulation platform runs scored, branching decision scenarios for many organizations at once. Shared infrastructure carries all of them, & hard tenant boundaries keep each organization’s data its own. Spend & audit trails stay separate the same way.

Business simulation was the first place we shipped this. The same design moves anywhere decisions carry consequences & a reviewer needs to see the reasoning behind the answer. Five sectors ask for the same engine with the domain swapped out. The generative AI development scope stays the same across all of them.

Sector The equivalent problem What changes in the build
Corporate learning & development Compliance & leadership training tests recall rather than decisions under pressure Swap the rubric for the competency model HR already runs
Healthcare clinical training Nurses & residents rehearse rare cases too dangerous to practice live Clinical review gates on content, with stricter audit under health-record rules
Financial services training Advisors learn suitability & disclosure from slides, then face real clients A scenario library tied to the regime, with scoring against conduct rules
Aviation & defense training Crews drill emergency procedures where a wrong call cascades fast Tighter scenario branching, with evaluation mapped to the certifying body’s framework
Government & public-sector training Caseworkers make judgment calls the agency must later defend Longer audit retention, with data residency inside the agency’s own boundary
Airline crew reviewing a simulator training session with an instructor at a debrief table
The same consequence engine fits sectors where crews rehearse high-stakes decisions before facing them live.

Porting needs a new scenario library & a rubric the sector trusts, set inside the buyer’s compliance regime. The agents, the isolation proof, the spend caps & the audit spine carry over unchanged.

Where this platform goes next

Next, this platform becomes our reference pattern for regulated multi-tenant AI products. The agents change with each domain, while the isolation proof & spend enforcement stay the same.

The audit spine carries over to every new domain as well. If your platform is one tenant away from needing all of this, the architecture in this write-up is where we would start.

You can read more about how we work on the About Brainy Neurals page.

Questions buyers actually ask

Buyers tend to ask these questions when they scope a multi-tenant AI platform.

How much does it cost to build an AI business simulation platform?

A production platform of this shape, with a multi-agent engine, multi-tenancy & compliance logging, is usually a multi-month build. Agent count & evaluation depth move the number, along with the tenancy model & the compliance regime. How much of an existing platform migrates or retires matters too. Sharing your constraints through the form gets you a range faster than a spreadsheet of assumptions.

Can AI reliably score student reasoning?

Yes, reliably enough for formative assessment, as long as the scoring is engineered. This platform scores against an explicit competency rubric with three strictness modes & tracks reasoning depth apart from correctness. Every evaluation is logged for professor review, & framework detection flags what students applied rather than merely mentioned. Professors keep final grading authority while the AI supplies evidence at a scale humans can’t match.

What makes an AI education platform FERPA-compliant?

FERPA compliance comes from architecture plus contract. On the architecture side you need access controls on student records & audit logging of reads, writes, deletes & admin actions. You also need a retention & deletion policy, & student data must stay out of model training. The contract side is the school-official designation with explicit data-use limits, as the US Department of Education’s student privacy guidance explains. EU deployments face a parallel review under GDPR & the EU AI Act, while Australian ones face the Privacy Act.

How do you test multi-tenant data isolation?

You test it adversarially, with tests written from the attacker’s side. A good suite has tenant A request tenant B’s objects by ID & cross scope through joins. It also probes the admin screens for scope leaks. Each test must fail before a fix ships & pass after it, & the suite runs on every deploy. A third-party audit found three such paths here, & the adversarial suite is why they stayed closed.

Shared schema or schema per tenant, which fits?

Shared schema with enforced tenant scoping fits many SaaS products until compliance pressure gets serious. It gives you one migration path & one connection pool at the lowest cost. It also leaves you one missed filter away from a leak, which is why adversarial testing matters so much there. Schema per tenant or database per tenant buys stronger boundaries at a real operating cost in migrations & backups. Our single-tenant to multi-tenant migration playbook walks through that decision.

How do you keep LLM costs predictable as adoption grows?

Start with attribution & add enforcement once spend is visible. Tag every model call with its tenant & its user at request time. Record the task too, then price each call against live provider tables in a real-time ledger. Next, give every tenant daily & monthly caps with alerts & a deliberate slowdown path when a cap hits. Multi-provider routing keeps unit costs inside chosen bounds, & costs stay predictable because the platform refuses overruns instead of reporting them later.

Why move from Node.js to Python for AI workloads?

The AI ecosystem centers on Python, from evaluation tools & data libraries to agent frameworks & the hiring pool around them. FastAPI with SQLAlchemy 2.0 gave this platform typed contracts & async speed close to the Node original, while opening the toolchain the roadmap needed. The migration cost is real, & six parity-gated milestones kept it manageable.

Tell us what you're building

Share your current stack & the compliance rules your buyers cite. The person who would architect your platform reads every message & replies.







    Services behind this case study

    One Brainy Neurals service line delivered this build from the agents down to the infrastructure.

    Generative AI development

    Multi-agent systems & evaluation layers built to run in production for many tenants.

    Similar case studies

    Similar case studies show the same patterns under different workloads.

    Real-time voice AI English tutor

    Sub-second conversational tutoring built to survive live classroom load.

    Multi-agent browser automation

    Where we first hardened the queue & fallback pattern behind this engine.

    AI tender document analysis

    The audit-first, provenance-heavy habit we carried into isolation proof.

    You can browse every build in Brainy Neurals case studies.

    Cite this case study

    Cite this case study with the reference below, which names the page & its permanent address.

    Brainy Neurals (2026). AI Business Simulation Platform for US Universities. Brainy Neurals case study. https://www.brainyneurals.com/case-studies/ai-business-simulation-platform

    Sources cited on this page