Home / Case studies / Automated Competitive Intelligence With AI for Diagnostics Makers
Case study · Life sciences · Generative AI
Automated Competitive Intelligence With AI for Diagnostics Makers
A US diagnostics & life sciences manufacturer needed automated competitive intelligence because analysts checked dozens of competitor & regulator sites by hand. Brainy Neurals built a collection platform that uses a language model to read each new article & file it under one category per company. Every week the pipeline runs on the client’s own cloud tenancy & replaces the rounds of opening sites one at a time. Analysts now search one archive of full article text, & adding a competitor takes one row in a source list.
Dozens
Sources read without hand checks
1 row
Adds a new competitor
Full text
Kept for later questions
Published October 2026
At a glance
What problem did this solve?
A diagnostics manufacturer tracked competitors by hand across dozens of newsrooms, trade titles & health agencies. Important signals arrived late, & some never arrived at all.
What did Brainy Neurals build?
Brainy Neurals built a competitive intelligence platform for a US diagnostics & life sciences manufacturer. It collects, reads, sorts & archives every article its tracked sources publish.
What changed after it went live?
Analysts stopped opening sources one at a time. New articles now arrive already read, already sorted into a category & already searchable next to everything published before them.
Who else could use this?
The same pattern fits any team that watches a named list of outside sources. Banks, manufacturers, carriers & retailers run the same problem every week.
Engagement facts
| Field | Detail |
|---|---|
| Industry | Healthcare & life sciences |
| Sub-vertical | Diagnostics & testing |
| Client type | US diagnostics & life sciences manufacturer |
| Engagement | Competitive intelligence platform |
| Timeline | Not disclosed |
| Capabilities | Generative AI, document AI, search |
| Delivery model | Project-based delivery |
Why did competitor news keep arriving late?
Competitor news arrived late because a US diagnostics & life sciences manufacturer tracked it by hand, with no automated competitive intelligence behind it. Analysts opened newsrooms, agency pages & publication feeds in turn.
Reading incoming text & filing it is intelligent document processing, & every step of it ran on people.
Where it broke
- A competitor posted a clearance on a Friday, & the team saw it the following Wednesday.
- One analyst filed a story under partnerships, & another filed the same story under product launches.
- Somebody asked what a competitor had announced in March, & the answer lived in old email threads.
- A trade publication changed its feed, stopped returning anything & went unnoticed for weeks.
- Two more competitors joined the watch list, & the weekly checking round grew again.
A clearance posted on a Friday reached the team the following Wednesday.
Information overload is an old problem for research teams. Work on it has long held that volume beyond a person’s capacity to process it lowers the quality of decisions[1].
What do teams try before automating?
Teams usually try four routes before they automate competitor tracking, & every one of them is reasonable. Each route holds in one setting & stops in another. The last row is the route Brainy Neurals took on this build.
| Approach | What it gets right | Where it stops | Who it still suits |
|---|---|---|---|
| Manual rounds | Judgment on every article | One analyst’s week | Two or three competitors |
| Keyword alerts | Free & instant | Noise, misses, no archive | Brand mentions only |
| Off-the-shelf suites | Coverage out of the box | Their sources, their categories | Standard sectors |
| Custom pipeline, our route | Your sources & categories | Weeks of crawler work | Many niche sources |
Each of the first three routes stops at a wall the fourth one clears.
A 2024 study of automated text annotation measured how steady model labels are. Lowering the model’s randomness raised agreement between runs from about 92% to about 98%[2].
How we designed the collection pipeline
Brainy Neurals designed the collection pipeline as one scheduled pass over a source list. Every stage writes its result before the next one starts.
A fetcher pulls new entries, & a crawler opens the ones nobody has seen yet. A language model reads each article & returns one category from a fixed list. A normalizer then lines up the company & source names.
Only the tracked sources & the category model sit outside the platform boundary.
Who owns the category list
The first decision was who owns the category list. Letting the model write its own labels was ruled out early, because a category that drifts can’t be counted.
What gets stored
The second decision was what gets stored. Brainy Neurals keeps the full article text instead of a summary, so a question nobody planned for still finds an answer.
Search runs over the stored text instead of anyone’s notes about it. Rushabh Shah reviewed every category decision against the way the client already reports. The model never invents a category & only picks one from a list the client owns.
The model never invents a category & only picks one from a list the client owns.
The same shape recurs across our LLM development work, with a different kind of text arriving at the front.
The technology stack we used
The technology stack had to survive a source that changes without warning. Brainy Neurals picked each layer for how it behaves when the input misbehaves. The same layers carry our knowledge base builds, with internal documents instead of news.
Eight layers fall into three groups, each picked for how it behaves when a source misbehaves.
| Group | Layer | What we used | Why | Ruled out |
|---|---|---|---|---|
| Collection | Source registry | Versioned list per company | One row adds a competitor | Hard-coded source URLs |
| Collection | Feed fetching | Bot-tolerant HTTP clients | Feeds return when requests fail | One generic HTTP client |
| Collection | Article retrieval | Headless browser crawler | Reads late-rendering pages | Raw HTML parsing alone |
| Reading & labeling | Language model | Small hosted model, gateway in front | Cheap per article, swappable | One large model everywhere |
| Reading & labeling | Prompt layer | Single label, fixed category list | Every article lands once | Free-text labels |
| Reading & labeling | Normalization | Rules over model output | Grouping survives messy names | Model-side name cleanup |
| Storage & delivery | Records & archive | Document store, archive, search index | History without re-crawling | Latest run only |
| Storage & delivery | Delivery | Scheduler, API & dashboard, separate | A stuck task spares the rest | One process doing everything |
How does one article get processed?
One article moves through six steps in the order the pipeline touches it, from the weekly wake-up to the search index.
- The scheduler wakes on its weekly slot & works down the source list, one tracked company at a time.
- Each feed is fetched with a client that survives bot checks, & entries the index has already seen are dropped.
- The crawler opens every new article in a headless browser, clears the consent screens & pulls out the body text.
- Headline & body go to the language model with the fixed category list, & the model returns exactly one label.
- Company & source names are matched against lookup tables, so one organization always groups under one name.
- The record lands in the database, the text lands in the monthly archive & the search index picks it up.
Retrieval steps up through three strategies, & a page with no body is recorded as a failure.
None of the six steps needs an analyst to open a source or file an article.
What broke & how we fixed it
Four things broke on this build, roughly in the order below. Each fix reads as obvious on paper, & none of them felt obvious at the time.
Blocked pages
The first crawl came back full of consent screens. Pages returned a 200 status code & no article, & the pipeline stored them as if they were stories.
The fix turned retrieval into a ladder of strategies that clears consent screens before reading the text. A page with no body now fails loudly instead of being stored quietly.
Inconsistent names
The same company arrived under four spellings across four feeds. Counts per competitor came out wrong, & nobody could tell by how much.
The fix sends company & source names through lookup tables the client can edit. One organization now keeps one name, however a feed spells it.
Silent feeds
One publication changed its feed structure & started returning nothing. The run finished, the dashboard looked healthy & a source had quietly gone dark.
The fix gives every source a health record, & a feed that returns nothing twice gets flagged on the dashboard. A broken source now shows up as a state instead of a silence.
Hedged labels
The model hedged more often than expected. Asked for one category, it sometimes returned two, & sometimes a sentence about why it couldn’t choose.
The fix gives the prompt a closed list & accepts one answer only. Anything else fails the record instead of slipping into the archive.
An audit of 14,000 web domains found automated access restrictions spreading fast across the sources worth reading[3]. Those are the weeks when clients bring in specialist engineers instead of learning the failure modes slowly.
Feed formats were standardized back in 1999, & the argument about versions never ended. That’s the kind of problem an AI proof of concept exists to find before anything is promised.
Every tracked source carries a state, so a feed that stops returning anything stops being invisible.
What changed once it went live?
After Brainy Neurals put the platform live, competitor news tracking moved from analysts opening sources one at a time to one scheduled run that fills one searchable archive. Each part of the work compares before & after go-live below.
One analyst opening sources in turn became one scheduled run filling one archive.
| What | Before | Now |
|---|---|---|
| Where articles are found | Analysts opening each source | Collected on a schedule |
| Who reads the full article | An analyst, when there was time | The pipeline, every time |
| How articles get categorized | Analyst judgment, varying by person | One fixed list |
| Historical coverage | Scattered across old digests | One searchable archive |
| Adding a competitor | More checking for somebody | A row in a source list |
| Knowing a source broke | Nobody knew | Tracked per source |
We haven’t published an hours-saved figure, though the client reports that the reading load dropped. A number we haven’t measured is a number we won’t print.
Day to day, an analyst opens one dashboard & sees what every tracked company published since the last run. Adding a competitor is a row in a list, & the weekly effort stays flat. Questions about last quarter get answered from the archive instead of from memory.
What the platform counts now
One label per article
The model returns exactly one label from the client’s fixed list, so every article lands in one place.
Three retrieval strategies
A plain request runs first, then a bot-tolerant client, then a headless browser. A page with no body is logged as a failure.
Three source health states
Every tracked source reads as healthy, stale or broken. Two empty runs raise a flag on the dashboard.
Want this pipeline over your own sources?
Send us the list of sources your team opens every week. We’ll reply with what a pipeline over them would look like & what it would take to build.
What is running today?
The automated competitive intelligence platform runs in production on the client’s own cloud tenancy, on its original weekly schedule. Each run collects, reads, sorts & archives whatever the tracked sources published that week.
Since handover, the client has added sources of its own, which takes no engineering time. The category list has changed twice, & the team that uses it made both changes. Analysts in life sciences now reach for the archive first, before a public search engine.
What would we do differently?
We would pin the categories first, plan more time for retrieval, watch every source from run one & review labels early.
A source that silently stops is worse than a source that loudly fails.
Health is recorded from the first run, so a gap raises a flag instead of passing unnoticed.
Pin the category list first
We wrote the prompt first & argued about categories afterwards. Every category change meant re-reading the articles already filed, which cost a week nobody planned.
Budget more time for retrieval
Most of the engineering went into getting the text at all. The model work took a few days, while the crawler work ran for weeks.
Instrument every source from run one
A source that silently stops is worse than a source that loudly fails. We now record health per source from the first run, & a gap raises a flag.
Review a sample of labels early
We shipped the classifier before anyone reviewed a sample of its labels. Half an hour of review in week one would have caught two categories that meant the same thing.
Where else does this pattern fit?
Automated competitive intelligence fits wherever a team has to know what a named set of outsiders published last week. The pattern collects articles from named sources & files each one under a fixed category after reading it.
| Industry | The equivalent problem | What changes in the build |
|---|---|---|
| Banking & finance | Regulator bulletins & competitor filings | Tighter categories, a label trail |
| Manufacturing | Supplier notices buried in trade press | Fewer sources, more weight on names |
| Logistics | Carrier advisories & port notices | Daily cadence, alerts over digests |
| Retail | Competitor launches on their own newsrooms | Product-level categories, images kept |
| Construction | Tender portals publishing bare listings | Document parsing behind each link |
Porting starts with a new source list & an agreed category list. Retrieval then gets one pass for the new sources, while the reading & filing stay exactly as they are.
The same collection pattern watching regulator bulletins instead of competitor newsrooms.
Questions about automated competitive intelligence
Buyers ask these six questions on almost every competitive intelligence call, & the full answers follow.
How it works
Can a language model categorize news articles reliably?
Yes, a language model categorizes news reliably when it picks from a closed list & the run uses low randomness. Published annotation studies put agreement between runs above 90% under those conditions. Open-ended labeling drifts, & drifted categories can’t be counted.
What happens when a source blocks automated access?
When a source blocks automated access, retrieval steps up through a ladder of strategies. A plain request runs first, then a client that survives bot checks, then a headless browser. A page that still returns no body gets logged as a failure instead of stored.
How is this different from a keyword alert?
A keyword alert only says that a phrase appeared somewhere, while this platform reads the whole article & gives it a business category. Each article is also tied to a normalized company name & kept searchable.
Time, cost & rollout
How long does it take to build competitive intelligence automation?
A working pipeline over a first set of sources usually takes a few weeks to build. Retrieval, the category list & the archive each need their own pass. A production version follows once the source health record holds steady.
How much does an automated competitive intelligence platform cost?
Cost depends on the number of sources, how hard they are to read & what already exists. Brainy Neurals scopes it from a short call & a look at your source list, then quotes a fixed price. An AI readiness assessment tells you first whether your sources can be read at all.
How many sources can one pipeline track?
Source count is a configuration question for one pipeline instead of an engineering one. Adding a competitor means adding rows to a list, & run time grows with article volume. Cost per run follows that volume, because every classification call records its own token count.
Which sources do you need to watch?
Name the newsrooms, feeds & agency pages your team checks by hand. One reply comes back from the person who would architect the pipeline, with no sales sequence behind it.
Services behind this case study
Six Brainy Neurals services sit behind this build, & each card below opens the page for that service.
Generative AI applications
Language model pipelines that read incoming text & return a structured answer every time.
Document AI
Intelligent document processing for unstructured text arriving from feeds, portals & inboxes.
RAG development
Searchable knowledge bases built over your own archive, so old coverage answers new questions.
AI agents & workflow automation
Scheduled workflow automation that runs the collection pass & hands the results to people.
Hire AI developers
Engineers who extend your team through the weeks when the sources fight back.
AI in healthcare
Clinical & life sciences systems built for teams working under regulatory scrutiny.
An AI proof of concept is the fastest way to test this against your own source list. AI consulting helps decide what to track before anything is built.
An AI readiness assessment shows whether those sources can be read at all. The industries hub shows where this pattern already runs.
Similar case studies
Three more Brainy Neurals builds run in real environments instead of in demos.
AI Diet Assistant for Gastroenterology
Clinical dietary guidance generated under review gates, live in a healthcare setting.
Personalised AI Meal Planning for Chronic Care
Structured meal plans built from messy personal health data, grounded & reviewable.
Overhead Line Geometry Measurement
Stereo cameras on a moving train measuring wire geometry, with inference on board.
Cite this case study
Shah, Rushabh. Automated Competitive Intelligence With AI for Diagnostics Makers. Brainy Neurals, September 2026. https://brainyneurals.com/case-studies/automated-competitive-intelligence-platform/
Sources cited on this page
- Eppler MJ, Mengis J. The Concept of Information Overload: A Review of Literature from Organization Science, Accounting, Marketing, MIS, and Related Disciplines. The Information Society. 2004;20(5):325-344. DOI 10.1080/01972240490507974.
- Alizadeh M, Kubli M, Samei Z, Dehghani S, Zahedivafa M, Bermeo JD, Korobeynikova M, Gilardi F. Open-Source LLMs for Text Annotation: A Practical Guide for Model Setting and Fine-Tuning. arXiv:2307.02179 [cs.CL]. 2024. DOI 10.48550/arXiv.2307.02179.
- Longpre S, Mahari R, Lee A, Lund C, Oderinwale H, Brannon W, et al. Consent in Crisis: The Rapid Decline of the AI Data Commons. arXiv:2407.14933 [cs.CL]. 2024. DOI 10.48550/arXiv.2407.14933.








