GenAI Observability for Assistant Usage & Cost in Diagnostics

Home / Case studies / GenAI Observability for Assistant Usage & Cost in Diagnostics

GenAI Observability for Assistant Usage & Cost in Diagnostics

A global molecular diagnostics enterprise lacked GenAI observability, so nobody could tell leadership how its internal AI assistants got used or what they cost. Brainy Neurals built a monitoring platform that uses shared metric definitions to compare usage, cost, feedback & access across every assistant. Read-only collectors pull the records each assistant already writes, so commercial & finance leaders read one dashboard instead of analyst spreadsheets. Analysts no longer assemble quarterly usage reports by hand, & leadership now reads adoption, cost, latency, feedback & access in one place.

  • #GenAIObservability
  • #GenerativeAI
  • #MolecularDiagnostics
  • #AICostMonitoring
  • #AIGovernance

Read-only

Source systems stay untouched

One definition

Numbers agree across teams

In one place

Where leadership reads spend

Mitesh

Published October 2026

At a glance

Brainy Neurals built one AI usage dashboard for a global molecular diagnostics enterprise, covering every internal AI assistant it runs.

What problem did this solve?

A global molecular diagnostics enterprise ran several internal GenAI assistants & couldn’t compare them. Usage, cost, feedback & access data each sat in a separate store.

What did Brainy Neurals build?

Brainy Neurals built an observability platform for a global molecular diagnostics enterprise. One dashboard now reports usage, cost, latency, feedback & access for every assistant.

What changed after it went live?

Analysts no longer assemble quarterly usage reports by hand. Leadership reads adoption & spend in one place, & access activity can be reviewed without opening each system.

Who else could use this?

The same observability layer fits any company running more than two internal AI assistants. Banks, insurers, manufacturers & logistics operators tend to reach this point after their second rollout.

Engagement factDetail
IndustryHealthcare & life sciences
Sub-verticalMolecular diagnostics
Client typeGlobal molecular diagnostics enterprise
EngagementObservability platform for AI assistants
TimelineNot disclosed
CapabilitiesGenerative AI, data analytics
Delivery modelProject-based delivery
TeamMitesh Patel & Rushabh Shah

Why couldn’t anyone answer basic usage questions?

Nobody at the enterprise could answer basic GenAI observability questions, because each internal AI assistant kept its own records. The company builds instruments for diagnostics laboratories, & its commercial & finance teams used the assistants to query business data & internal sources.

When leadership asked what the assistants cost, the answer took two weeks of spreadsheet work. Each assistant came from a separate LLM development effort, with its own store & its own definitions.

Where it broke

  • Quarterly reviews began with an analyst exporting raw records into a sheet.
  • Session counts differed by application, so two decks disagreed.
  • Cost questions arrived after the invoice, since nobody totalled tokens until asked.
  • Feedback ratings stayed beside their own application, out of sight when comparing assistants.
  • Access reviews meant opening each system & reading logs that didn’t match.

None of this was unusual for a company whose AI estate kept growing. A vision paper on production machine learning observability describes failures after deployment as silent, arriving with no feedback to catch them[1].

Analyst desk covered in printed AI usage reports being reconciled by hand at the end of a day
Before the platform, a quarterly usage review took two weeks to assemble by hand.

Why don’t the obvious fixes hold up?

Four routes usually get tried before a shared observability layer, & each one is reasonable for a while.

One reporting need branching into per-app dashboards whose totals disagree, spreadsheets that go stale, a provider console that shows cost only, and a shared layer that gives one view. per-appspreadsheetconsoleshared layer totals disagreegoes stalecost only one view

Three common routes each stop at a wall that the shared layer avoids.

ApproachWhat it gets rightWhere it stopsWho it still suits
Per-app dashboardsOwned by its own teamDefinitions drift, totals disagreeOne assistant, one owner
Spreadsheet reportingNo new system, starts todayPulls age, errors compoundA one-off board question
Provider consoleAccurate for that provider’s spendBlind to feedback & accessSingle-provider cost estimates
Shared layer, our routeEvery application measured alikeWeeks of reconciliation up frontA growing family of assistants

Spreadsheet reporting carries its own risk, too. Four field studies covering more than 300 working spreadsheets each found error rates the reviewer called unacceptable[2].

How does the GenAI observability layer work?

Brainy Neurals’ observability layer reads records & never writes them. We connected it read-only to each assistant’s own store. It pulls records that already exist & reshapes them into one layout before anything is counted.

Nothing on this platform calls a model, because it only reads what the models already left behind.

Architecture of the GenAI observability platform. Three assistants feed a read-only pull, a contract check that holds failures in quarantine, one metric layer, cached tables and dashboard pages behind sign-on and roles, all inside the client tenant. The model provider sits outside and is never called. SOURCESPIPELINEACCESS tenant held backrole gate Assistant AAssistant BAssistant CRead-only pullContract checkQuarantineMetric layerCached tablesDashboardSign-on & roles Model providernot called GenAI observability platform, stacked. Assistants A, B and C feed a read-only pull, a contract check with quarantine, one metric layer, cached tables and dashboard pages behind sign-on and roles, inside the tenant. The model provider is never called. Assistant A,Assistant B, CRead-only pullContract checkQuarantineheld back Metric layerCached tablesDashboardSign-on& rolesrole gate tenant Model providernot called

Each record is read once, measured once & served from a cache, with no model call anywhere on the path.

Nothing on this platform calls a model, because it only reads what the models already left behind.

Our first decision was to settle the metric definitions before writing a single connector. A session, an active user, a token cost & a satisfied rating each got one definition, agreed once & never redefined inside one application.

The second decision treated every source as untrusted input. Field names differ between applications, & they change between releases of the same application. Each record passes a shape check on the way in. Anything that fails lands in a quarantine lane instead of averaging quietly into a director’s number.

Mitesh Patel set those metric definitions & reviewed every record contract before a connector was written. Rushabh Shah built the collectors & contract checks, then wired the metric layer to the cache.

Settling definitions first is AI consulting work more than engineering, & skipping it is why per-app dashboards disagree.

The technology stack we used

The technology stack had to survive an audit & a schema change on the same day. We chose each layer for traceability, because a metric nobody can explain is one nobody acts on. The same layers sit under our AI copilot builds, with different records underneath.

Data & collection

Source store. The existing document database already held the records, so no second production copy was made.

Collectors. One read-only collector per application means no write path, & nothing runs inside an assistant.

Shape check. A contract per source makes drift fail loudly, since field names are never trusted.

Metrics & analysis

Metric layer. One definition per measure makes applications comparable, so we ruled out per-app metric code.

Caching. Pre-computed tables for each page keep production untouched, instead of one query per filter.

Application & access

Interface. A Python web app with charts built in is quick to extend & needs no per-seat licence.

Sign-on & roles. Enterprise single sign-on means access follows the company directory instead of a separate user table.

Deployment. Containers with environment secrets give the same build everywhere, with no hand-configured servers.

How does one record reach the dashboard?

One usage record passes through six stages between an assistant & a reader, & only one stage touches the source store.

  1. Each assistant writes its own records where it already writes them, & nothing about that changes.
  2. A read-only collector pulls new records on a schedule, with one connector per application.
  3. A shape check compares each record against its contract, & any failure goes to quarantine with a reason.
  4. The metric layer computes every measure from one definition, so any two applications can be compared.
  5. Results land in a cached table, which the dashboard reads instead of the source store.
  6. A signed-in director opens a page, a role check decides which applications appear, & the filters run.
Six steps of one usage record: record written, collector pulls across the only crossing into the platform, contract check with a quarantine for failures, metric computed, cache written, page served behind a role gate, with no live query back to the source. SOURCE SIDEPLATFORM SIDE no live query 1Record written2Collector pulls3Contract check4Metric computed5Cache written6Page served only crossingheld Quarantine role gate Six steps of one usage record, stacked: record written, collector pulls across the only crossing, contract check with quarantine, metric computed, cache written, page served behind a role gate, with no live query back. SOURCE SIDE onlycrossing PLATFORMSIDE 1Record written2Collector pulls3Contract check4Metric computed5Cache written6Page served Quarantineheld role gate no live query

A record that fails its contract stops at step three & is counted, so it never vanishes quietly.

What broke & how we fixed it

Three problems broke the build in the field, in this order, & each one needed a different fix.

Overlapping printed report pages pinned across a wall at night, each with a different layout
No two applications printed a report the same way, which is why two decks never agreed.

Records that changed shape

The first collector ran clean for a week, then returned half the sessions. One application had renamed a field in its own release, & nothing noticed, because the records carried no declared schema.

We wrote a contract for every source that names the fields each record must carry. Failures go to quarantine with a reason, & the front page prints how many were held back.

Counts that disagreed

Two directors compared the same month & got different active-user counts. Our definition needed a user identifier on every event, but one application wrote it only at login, so its curve looked like collapsing adoption.

We changed the definition instead of the data, which cost less. An active user now needs one attributable event that day, each application declares which events qualify, & the low curve straightened within a release.

Pages that stalled

Every filter change re-queried the store behind the page, once per chart. A page with six charts took so long that people stopped moving filters.

We moved the heavy aggregates onto a schedule & pointed the pages at the cache. Filters now run against thousands of rows instead of the whole history.

A contract check feeding a dashboard panel of usage, cost and feedback, with a held-back record count on the front page. contract usagecostfeedbackheld back

The held-back count sits on the front page, because a pipeline that hides its rejects gets audited the hard way.

The word quarantine comes from the forty days Venetian ships waited offshore, which the engineers who named the lane didn’t know.

A study of ten applications on schema-free stores found the schema changing in every one[3]. Weeks like these are when teams bring in specialist engineers. Surfacing that work before a platform is promised is what an AI proof of concept is for.

What changed after go-live?

Go-live changed how the diagnostics enterprise reads its whole AI estate, starting with five working habits.

DimensionBeforeNow
Assembling a usage reportManual export, then spreadsheet workThe page is already built
Where a number comes fromEach application’s own storeOne definition, applied everywhere
Noticing a cost riseOnly after the invoice arrivedOn the day it moves
Reviewing who has accessEach system opened in turnOne audit view across all
Adding a new assistantA new reporting processA connector & a contract

Each assistant now reaches the platform through one connector & one contract that cover six kinds of record. Analysts maintain the metric definitions instead of assembling quarterly exports, & commercial & finance leaders read adoption & cost from one dashboard.

We haven’t published any outcome figures for this build. The client reports that each of those numbers moved the right way, & a number we haven’t measured is one we won’t print.

A director opens one page instead of asking an analyst for a pull. Because every metric comes from one definition, two assistants can be compared without a footnote. Onboarding the next assistant into GenAI observability costs a connector instead of a habit.

The platform runs in production today across the client’s family of internal AI assistants. Since handover, new assistants have joined through the same contract & connector path, & the platform still calls no model. A business in healthcare that expects audits values that restraint.

Two colleagues reading a single printed AI usage summary at a standing table beside a lab window
One printed page now carries what five separate exports used to, & the binders behind it stayed shut.

Running more than one internal AI assistant?

Tell us which usage or cost numbers you can’t see today. We’ll reply with how an observability layer would read your own assistants’ records.

What would we do differently next time?

Four calls on this build would change if Brainy Neurals ran it again, & each one is cheap to copy.

Agree definitions with the people who present them

Next time, we’d settle every metric definition with the directors who present the numbers, before the first release.

On this build we settled them with the engineers who own each application. The directors joined two releases later, & one definition changed on the spot.

Write the record contract before the first connector

We’d write each source’s record contract first & build its collector against that contract.

Here we built collectors against the data we could see that week, & that data moved underneath us.

Assume the cache from the start

Every page would read pre-computed tables from the first day.

We queried the store live because the first application was small, & that choice lasted exactly as long as the third application.

Put the quarantine count on the front page

The held-back record count would sit where every reader sees it.

We first logged rejected records where only engineers looked. A pipeline that hides what it rejected is one nobody audits, & audits were the reason for this build.

Where else does this pattern fit?

GenAI observability measures how AI applications are used, what they cost & how they perform, using records they already write. The pattern belongs wherever more than one AI assistant runs inside a company.

AI assistants in a bank’s retail, lending and treasury units report into one observability panel, answering who used what. RetailLendingTreasury who used what

The same pattern over a bank’s assistants, where the regulator asks first instead of the finance team.

Banking & insurance

Teams in banking reach this point once compliance asks who used what across three business units. Retention gets longer, & the rest of the build stays the same.

Manufacturing

In manufacturing the question arrives when two plant copilots post different cost lines. Site becomes a dimension in every metric.

Logistics

Dispatch & service assistants often draw on one budget. Reporting moves from per day to per shift, because that’s how the work runs.

Retail

Store assistants see seasonal traffic swings that monthly averages hide. Forecasting matters more than history in that build.

Professional services

Research assistants get billed to client matters. Every event needs a matter code before the numbers mean anything.

Construction

Bid & site assistants work across shared projects. Access splits per joint venture in that build.

Porting the pattern to a new company takes a connector & contract per source, plus one argument about sessions.

Questions buyers usually ask

Buyers usually ask these six questions about AI assistant monitoring before they scope a build.

What is GenAI observability?

GenAI observability is the practice of measuring how generative AI applications behave in production. It covers usage, cost, latency, feedback & who had access, collected from the records those applications already write.

How is AI observability different from app monitoring?

Application monitoring answers whether a service is up & how fast it replied. Observability also answers who used it, what the tokens cost & whether the access was appropriate. It shows whether the person who asked rated the answer useful.

What should an AI usage dashboard measure?

An AI usage dashboard should start with active users, sessions, token cost per application, response latency & feedback ratings. Add access & audit events if you work under a regulator. Each measure needs one definition applied across every application.

Should we build this in-house or buy a tool?

Buy a tool when your assistants all sit on one framework that a vendor already parses. Build when the records live in your own stores & the data can’t leave your tenant. Build too when the definitions are contested inside the company & need settling first.

How long does an AI usage dashboard take to build?

The first application takes the longest, because the metric definitions get argued once & settled once. Each assistant after that is a connector, a contract & a mapping, which lands in weeks instead of quarters.

How much does an observability build cost?

Cost tracks the number of applications, how far their record shapes have drifted apart & whether single sign-on already exists. Brainy Neurals scopes it from a short call, then quotes a fixed price. An AI readiness assessment tells you first whether your records can carry the metrics you want.







    Services behind this case study

    Six Brainy Neurals services carried this build from the first definition argument to production.

    Generative AI applications

    Internal assistants that answer from your own data, built to be measured from day one.

    AI consulting

    Metric definitions & the governance model, settled before anybody writes a connector.

    AI agents & copilots

    Copilots that log what they did, so usage & cost can be read later.

    RAG development

    Retrieval over internal knowledge, with the retrieval itself instrumented.

    Hire AI developers

    Data & platform engineers who extend your team while the schemas keep moving.

    AI in healthcare

    Systems for life-sciences businesses that expect an auditor to ask.

    An AI proof of concept tests this against two of your own applications before anything larger gets scoped. Our engagement models cover how a platform like this gets staffed & handed over. An AI readiness assessment tells you whether your records can carry the metrics you want. The industries hub shows where the pattern already runs.

    Similar case studies

    Another Brainy Neurals build shipped generative AI into a working healthcare setting under review gates.

    AI Diet Assistant for Gastroenterology

    Clinical dietary guidance generated under review gates, live in a healthcare setting.

    Cite this case study

    Patel, Mitesh & Shah, Rushabh. GenAI Observability for Assistant Usage & Cost in Diagnostics. Brainy Neurals, September 2026. https://brainyneurals.com/case-studies/genai-observability-dashboard/

    Sources cited on this page

    1. Shankar S, Parameswaran AG. Towards Observability for Production Machine Learning Pipelines. Proceedings of the VLDB Endowment, 2022, 15(13), 4015-4022. DOI 10.14778/3565838.3565853. arXiv:2108.13557.
    2. Panko RR. What We Know About Spreadsheet Errors. Journal of Organizational and End User Computing, 1998, 10(2), 15-21. DOI 10.4018/joeuc.1998040102.
    3. Scherzinger S, Sidortschuck S. An Empirical Study on the Design and Evolution of NoSQL Database Schemas. Conceptual Modeling, ER 2020, LNCS volume 12400, pages 441-455. DOI 10.1007/978-3-030-62522-1_33. arXiv:2003.00054.