Agentic Text-to-SQL Assistant for a US Banking Group

Home / Case studies / Agentic Text-to-SQL Assistant for a US Banking Group

Case study · Banking & finance · Agentic AI

Agentic Text-to-SQL Assistant for a US Banking Group

A US banking group needed immediate answers from its customer & transaction databases, yet every question waited on engineers translating it into SQL. Brainy Neurals built an agentic text-to-SQL assistant that uses open-source language models to turn plain English questions into checked database queries. The assistant runs inside the bank’s own network, where business users ask in a web chat instead of filing a ticket. Answers that took up to a full working day now arrive in seconds, & no customer data leaves the building.

  • #TextToSQL
  • #AgenticAI
  • #GenerativeAI
  • #RetailBanking
  • #NaturalLanguageQueries

Seconds

Time to an answer

In the network

Where every model runs

Any business user

Who can ask the warehouse

Mitesh

Published October 2026

At a glance

What problem did this solve?

Business leaders at a US banking group couldn’t query customer & loan data on their own. Every question went through the data teams, & an urgent answer could take a full day.

What did Brainy Neurals build?

Brainy Neurals built an agentic text-to-SQL assistant for a US banking group. It turns plain English questions into SQL & runs them on the warehouse, then sends back rows with a short summary.

What changed after it went live?

Leaders now ask the bank’s warehouse questions in plain English & get answers in seconds. Answers that once took a full day arrive while the meeting is still going, & the data teams are back to engineering.

Who else could use this?

Natural language database queries suit any organisation where decision makers wait on technical teams for numbers. Insurers, hospitals, logistics firms, retailers & manufacturers keep the same kind of locked warehouse.

The engagement in brief

  • Industry · Banking & financial services
  • Client · A US banking group, retail & commercial
  • Engagement · Agentic AI data assistant
  • Timeline · Not disclosed
  • Capabilities · Generative AI & agentic AI
  • Delivery · Project-based

Why did every bank report take a day?

Every report took a day because each question had to become SQL, & only the data teams could write it. The US banking group wanted an agentic text-to-SQL assistant so leaders could ask the warehouse directly.

The bank runs its decisions on a large relational data warehouse. Inside it sit the customer profiles & account activity that bankers watch, along with loan records & insurance policies. Leaders asked questions like which customers hold active policies every day.

SQL is the query language a database answers to, & only the data teams wrote it. Leaders wanted a conversational AI layer over the warehouse that anyone could use.

Hands leafing through a printed banking report, the manual data request routine that came before natural language database queries
Before the assistant, a question left as a ticket & came back days later on paper.
  • A question left a business team as a ticket & came back days later as a spreadsheet.
  • Urgent requests queued behind routine ones, because one team wrote all the SQL.
  • Each follow-up question restarted the cycle, since a static export can’t answer the next one.
  • The data teams spent their days turning English into queries instead of engineering.

People have wanted to talk to databases for a few decades now. A 2019 survey compared 24 natural language interfaces for databases & found that even the strongest leaned on hand-built rules[1].

What do banks try before a text-to-SQL assistant?

Banks usually try four routes before they build a natural language query assistant, & each route makes sense for some teams.

Approach Strength Where it stops Still suits
BI dashboards Answers predicted questions New questions need new dashboards Stable, recurring metrics
SQL training for staff Needs no new software Fails on thousands of cryptic columns Analysts close to the data
Cloud AI copilots Strong SQL generation Customer data leaves the network Public or synthetic datasets
On-premise agent, our route Plain English on live data Weeks of schema & guardrail work Regulated data, non-technical users

On a 2023 benchmark of real enterprise databases, the strongest model reached 40% execution accuracy while humans scored 93%[2]. That gap is the reason the last row of the table exists.

How did we design the agentic text-to-SQL assistant?

We designed the query assistant as an agent that acts on the database & checks its own work. The agent reads the intent behind a question, then pulls the matching slice of schema from a vector database.

A language model hosted on the bank’s own servers then writes the SQL. The model never sees the whole warehouse, & the warehouse never sees the internet.

Mitesh Patel, our AI architect, set the agent architecture & checked every query-safety decision against the bank’s own data access rules.

Where the models live

We ruled out cloud model APIs from the start, because regulated customer data crossing that line would never clear review. Open-source models run instead on hardware the bank already owns & controls.

What the agent may do

We also ruled out a tool that only drafts SQL for someone else to run, since the person asking can’t judge a query. The agent runs each query with read-only credentials & checks what comes back. When the database reports an error, the agent tries again.

Architecture of the agentic text-to-SQL assistant. A business user asks in a chat interface. Inside the bank network, the agent service retrieves tables from a schema index, asks local open-source models to write SQL & runs it read-only on the data warehouse. Errors loop back to the models. Cloud AI sits outside with no connection. BANK NETWORK BUSINESS USER asks in English CHAT INTERFACE web chat AGENT SERVICE plans & checks SCHEMA INDEX finds the tables LOCAL MODELS write the SQL DATA WAREHOUSE runs read-only CLOUD AI no connection Architecture of the agentic text-to-SQL assistant, stacked. A business user asks in a chat interface. Inside the bank network the agent service uses the schema index, local models & the data warehouse. Cloud AI sits outside with no connection. BUSINESS USER asks in English BANK NETWORK CHAT INTERFACE web chat AGENT SERVICE plans & checks SCHEMA INDEX finds the tables LOCAL MODELS write the SQL DATA WAREHOUSE runs read-only CLOUD AI no connection

Every step from question to answer runs inside the bank’s network, & no call ever leaves it.

What technology stack did we use?

Every layer of the stack had to pass one test first, which was whether customer data could ever leave through it. The same layers show up across most of our generative AI development work, with a different database underneath.

Data & retrieval

A read-only connection to the existing warehouse means a wrong query can’t write anything. We ruled out broad service accounts for the same reason.

A vector database of plain-language column descriptions finds tables by their meaning. We ruled out putting the whole schema into every prompt.

Model & reasoning

Open-source language models run on the bank’s servers, so the data stays inside. We ruled out cloud model APIs for that reason.

A Python agent loop runs each query & checks the result, so errors turn into feedback. We ruled out one-shot generation, where the model gets a single try.

Each step gets its own prompt template, because compact models need a tight brief. We ruled out one generic prompt for every step.

Safety & delivery

A read-only role with row limits & timeouts protects the production warehouse. We ruled out simply trusting the model.

Users ask in a web chat with suggested follow-ups, which meets them where their questions start. We ruled out building yet another dashboard.

How does one question become an answer?

Each question passes through five steps, in the order the assistant actually handles them.

  1. Ask in plain English. A business user types a question in plain English, & the agent works out what it means.
  2. Find the schema. The agent turns the question into a search & pulls the matching tables & columns from the vector index.
  3. Write the SQL. A local language model writes SQL against that slice of schema, following a prompt template made for this step.
  4. Run it read-only. The agent runs the query with read-only credentials, & any database error goes straight back to the model to fix.
  5. Summarise & suggest. Results come back as a table with a plain-language summary, plus follow-up questions ready to ask next.

Customer data stays inside the bank’s network through all five steps. The loop in step four is what separates a demo from something a bank will sign off.

What nearly stopped the build?

Three problems nearly stopped the build, & they turned up in roughly this order. Each card below pairs what broke with how we fixed it.

Data centre server racks holding the bank's enterprise data warehouse behind the text-to-SQL assistant
The warehouse & the models share one building, & nothing leaves it.

Thousands of cryptic columns

Column names like CUST_TXN_DTL meant the model either saw too little context or drowned in it.

We stopped showing the model the schema & started retrieving it instead. Every table & column got a plain-language description in the vector index, so a question pulls back only the few tables that matter.

SQL that failed two ways

Some queries broke on syntax, which the database at least reported. Others ran cleanly but answered a slightly different question, & nothing flagged that at all.

We made the database part of the loop, with every query running read-only & each error going straight back for another try. Summaries must quote the returned rows, never the model’s memory, which catches the silent kind.

Models smaller than the headlines

The models allowed inside the bank were smaller than the ones making the news. What a big cloud model got right first time, a local one reached after retries.

We stopped asking one call to do everything. Separate prompt templates now handle each job, from writing a query to summarising its rows, each tuned to what a compact model does well.

Published research already pointed the way out. A 2023 study showed a model repairs its own failed queries when it sees the execution result, lifting hardest-problem accuracy by 9%[3].

The cryptic column names are inherited, by the way. Length limits in the mainframe era taught database designers to abbreviate, & the habit outlived the mainframes.

These are the weeks when banks start asking about specialist engineers instead of learning the hard way. They are also the weeks an AI proof of concept exists for, before anyone promises leadership a chatbot.

What changed once the assistant went live?

Once the assistant went live, answers that took hours, sometimes a full working day, started coming back in seconds.

Seconds is the client’s own word, & Brainy Neurals hasn’t measured a latency figure behind it. We won’t print a number we haven’t measured, & Ronak Patel wrote this case study from the build notes & the client’s own account.

What Before Now
Who runs a query Data teams, on request Any authorised business user
Time to an answer Hours, sometimes a full day Seconds, inside the chat
Follow-up questions A new ticket & a new wait Asked in the same conversation
The data teams’ day Turning questions into SQL Engineering
AI services outside the bank None permitted None needed

Day to day, the change shows up in ordinary meetings. A question raised at ten gets its answer at ten, while the decision is still open. Since the models run on the bank’s own hardware, there’s no per-question API fee, so exploring costs nothing extra.

The assistant runs in production today against the warehouse the data teams once served by ticket. Senior managers & business teams use it directly, suggested follow-ups included.

The data teams still own the schema descriptions & update them as tables change, which is now most of the upkeep. New tables join the vector index in the same week they appear, & no customer data ever leaves the building. Keeping that boundary is what got this AI in banking & finance project approved.

Want this assistant on your own schema?

Tell us which questions your team waits on, & we’ll show how an agentic assistant would answer them inside your network. You can also start with a short AI readiness check.

What would we do differently next time?

Four lessons came out of this build, sorted here into what we’d repeat & what we’d avoid.

What we would repeat

  • Describe the schema first, because the plain-language data dictionary moved accuracy more than swapping models did.
  • Collect the real questions early, since executives asked shorter, vaguer questions than we imagined.
  • Put the effort into the loop, where a model that reads its own error message beats a bigger model that can’t.
  • Show the query, because people trusted the assistant once they could open the SQL & rows behind every summary.

What we would avoid

  • Building the data dictionary second, which is what we did, & next time it comes first.
  • Designing for imagined questions, because the real ones arrived full of shorthand & the prompts had to change.
  • Swapping models to chase accuracy, since most of our gains came from the loop & almost none from model swaps.
  • Leading with accuracy claims, because one expandable panel did more for adoption than any number.

Where else does this pattern fit?

Natural language database queries turn plain questions into SQL & return short answers wherever decision makers wait on technical teams.

Insurance claims desks

Adjusters wait on report teams for claims data today. Porting needs a glossary of policy terms mapped to the right tables.

Healthcare operations teams

Managers ask about beds & waiting times, along with readmissions. The build adds de-identification rules with tighter access scoping.

Logistics control rooms

Operations teams ask where shipments sit right now. The main change is time-window phrasing over tables that change fast.

Retail merchandising teams

Merchandisers query sales without waiting for an analyst. Product hierarchies & seasonal vocabulary go into the schema descriptions.

Manufacturing plant floors

Plant managers pull quality & downtime numbers for themselves. Shop-floor terms get mapped onto the sensor tables.

The vocabulary changes from one industry to the next, while the loop underneath stays the same. Porting takes a described schema & a fresh vector index, plus question tuning with the people who will actually ask.

What do buyers ask about text-to-SQL assistants?

Can an AI chatbot query a SQL database directly?

An AI chatbot can query a SQL database directly when it is built as an agentic text-to-SQL assistant. It turns a plain English question into SQL & returns the rows with a short summary. Retrieval & guardrails around the model keep it reliable on enterprise schemas.

Why can’t a general chatbot answer from our database?

A general chatbot has never seen your schema, so it invents table names that sound right & runs nothing. An assistant wired to your database retrieves the real schema & runs queries read-only, then shows you its rows.

Is it safe to connect a language model to a production database?

Connecting a language model to a production database is safe when it gets read-only credentials with row limits & timeouts. A bad query then wastes seconds & changes nothing. Keeping the models on your own servers keeps the data off the internet.

How long does a natural language database assistant take to build?

An AI proof of concept on your schema usually takes a few weeks. Schema descriptions & retrieval tuning each need a pass, & so does the correction loop. Production hardening follows once real users have asked real questions.

How much does a text-to-SQL chatbot cost?

The cost depends on the size of your schema & the variety of questions, plus what already exists in your stack. Brainy Neurals scopes it from a short conversation & a look at the database, then quotes a fixed price. An AI readiness assessment shows first whether your data layer is ready.

Can this run without any cloud AI service?

A natural language database assistant can run entirely on your own servers, with no cloud AI service at all. Open-source models now handle the pattern well, & hosting them yourself means no per-query fees & no customer data leaving your network. Inside a bank, that usually settles it.

Tell us what your team needs to ask

Send the questions your leaders wait on. The person who would architect the assistant will reply with how it could answer them inside your network.







    Services behind this case study

    AI agent development

    Agentic assistants that act on your systems & check their own output before they answer.

    Generative AI applications

    Language model systems built for enterprise data, from chat assistants to summary pipelines.

    RAG development services

    Retrieval pipelines & vector databases that ground every model answer in your own data.

    Document AI services

    The same extraction discipline, applied to the statements & forms that banking runs on.

    Hire AI developers

    LLM & data engineers who join your team through the weeks heavy with schema work.

    AI in banking & finance

    Assistants & document pipelines for financial data that can’t leave your network.

    An AI proof of concept is the quickest way to test natural language database queries on your own schema. AI consulting helps pick the model & the network boundary first. An AI readiness assessment checks your data layer, & the industries hub shows where the pattern runs.

    Similar case studies

    AI Diet Assistant for Gastroenterology

    Clinical dietary guidance behind review gates, live in a healthcare setting. Like this build, it grounds generative AI answers with human checks & runs in production.

    Cite this case study

    Patel, Ronak & Patel, Mitesh. Agentic Text-to-SQL Assistant for a US Banking Group. Brainy Neurals, September 2026. https://brainyneurals.com/case-studies/natural-language-database-queries/

    Sources cited on this page

    1. Affolter K, Stockinger K, Bernstein A. A comparative survey of recent natural language interfaces for databases. The VLDB Journal, 2019, volume 28, issue 5, pages 793 to 819. doi.org/10.1007/s00778-019-00567-8
    2. Li J, et al. Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs. NeurIPS 2023 Datasets and Benchmarks. arxiv.org/abs/2305.03111
    3. Chen X, Lin M, Schaerli N, Zhou D. Teaching Large Language Models to Self-Debug. 2023. arxiv.org/abs/2304.05128