Home / Case studies / AI DICOM De-Identification & Organ Masks in Medical Imaging
Case study · Healthcare · Computer vision & document AI
AI DICOM De-Identification & Organ Masks in Medical Imaging
A US medical imaging organization needed DICOM de-identification on every CT & MRI study it shares with researchers, & staff handled each one manually. Brainy Neurals built a one-pass pipeline that uses a computer vision model to outline organs while it strips patient identifiers. The pipeline runs inside the client’s own network, reads each study once & replaced the manual header check before every handover. Each study now leaves with its patient identifiers removed & its organ masks attached as a standard DICOM-SEG file.
Inside the network
Studies never leave site
Same read
Scrub & segment together
DICOM-SEG
Masks open in current viewers
Published October 2026
At a glance
What problem did this solve?
A US medical imaging organization removed patient identifiers from CT & MRI studies by hand. That scrubbing held up every research handover, & organ segmentation then started from scratch.
What did Brainy Neurals build?
Brainy Neurals built an automated DICOM de-identification & segmentation pipeline that runs inside the client’s own network. One pass scrubs the header against the DICOM confidentiality profile, & a computer vision model writes the organ masks as DICOM-SEG.
What changed after it went live?
Studies now leave the department with identifiers stripped & organ masks attached. Nobody retypes a header, & segmentation no longer waits for a specialist.
Who else could use this?
Any group that moves imaging studies between systems, sites or partners can use single-pass de-identification. Hospitals, research networks, imaging vendors & trial sponsors all fit.
Engagement facts
| Fact | Detail |
|---|---|
| Industry | Healthcare, medical imaging |
| Client type | US medical imaging organization |
| Engagement | Imaging compliance platform |
| Timeline | Not disclosed |
| Capabilities | Computer vision, document AI |
| Delivery model | Project-based delivery |
Why does DICOM de-identification go wrong?
DICOM de-identification goes wrong when people do it by hand on top of a full clinical workload. A US medical imaging organization sends CT & MRI studies to researchers, vendors & partner sites, & every study has to lose its patient identity before it travels.
Scrubbing a study is document AI work, & most of it happens in the header. A DICOM header is the block of labelled fields, called tags, stored with every scan. Headers run long & drift into free text, & each scanner writes them its own way.
Field by field
A technician read each header & typed over any name they found.
One miss, one rework
A single name left in a description field sent a whole handover back for review.
Two queues
Segmentation waited until the scrubbing was signed off, so the two jobs never overlapped.
Slice by slice
A specialist drew organ outlines by hand while real cases waited in the queue.
Every scanner differs
Each machine wrote its header its own way, so last month’s rule failed this week.
Ten toolkits, one safe default
A 2015 evaluation tested 10 free DICOM de-identification toolkits for this weakness. Only one removed every required element on its default settings.[1]
Why do common approaches stall?
Common approaches stall because each of the four routes to scrubbing a study only fits some settings. By hand, tiredness decides what gets missed, & the other three stop at a wall the fourth one doesn’t meet.
| Approach | Strength | Where it stops | Still suits |
|---|---|---|---|
| Scrub it by hand | Human judgment on every field | One study at a time | Rare, small handovers |
| Off-the-shelf anonymizer | Covers the obvious tags fast | Misses private tags & free text | Single-scanner archives |
| Delete everything optional | Nothing identifying survives export | Dates & links break research value | One-off dataset drops |
| One automated pass (our route) | Identity & masks handled together | Weeks of profile & validation work | Studies moving in volume |
A 2015 review notes that deleting all private elements removes useful scientific data along with the identifiers.[2]
What Brainy Neurals built
Brainy Neurals built a one-pass imaging pipeline for a US medical imaging organization, so each study is read once & written once. The header work & the organ segmentation run over the same loaded volume on the client’s own network. Scrubbing & segmentation stopped being two queues waiting on each other.
One pass over one loaded volume, & nothing crosses the client’s network edge.
The first decision was where the rules come from, & we started from the confidentiality profile the DICOM standard publishes. A hand-kept tag list only knows the tags it has already met. We map identifiers & dates rather than deleting them. A study that loses its dates loses the reason somebody wanted it, so that choice was deliberate.
A study that loses its dates loses the reason somebody wanted it, so that choice was deliberate.
The second decision was what comes out the other end. Masks are written as DICOM-SEG against the de-identified study, so they open in the viewers the client already runs. We turned down a research-format side file, because a side file gets separated from its study within a week.
Teaching a model to outline organs is ordinary computer vision development, & everything around the model is what makes its output usable.
The stack we used
The stack runs inside the client’s own environment, because none of this data may leave it. That ruled out a hosted processing service before we chose anything else. We picked each remaining layer for what it costs to run & maintain, which is technology selection work done before any code.
DICOM read & written natively
The study keeps its structure from intake to export. We ruled out converting to a research format.
The standard’s confidentiality profile
The DICOM standard already names what carries identity. We ruled out a hand-kept deny-list.
Pseudonyms & shifted dates
Studies from one patient still group after the names go. We ruled out fresh random values.
Medical imaging deep-learning framework
It’s built for volumetric data, which is what CT & MRI scans produce. We ruled out a general vision library.
DICOM-SEG against the source series
Masks open in the viewers the client already runs. We ruled out a mask file sitting beside the study.
Python services with a watched intake
Python keeps one language across the whole toolchain. We ruled out a workflow product that needs a license.
How does one study move through the pipeline?
One study moves through six steps in the order the pipeline sees it, & a person appears in none of them.
Check the study at intake
A study lands in the watched intake, & the pipeline checks it’s complete before touching anything.
Read the full header
The pipeline reads the full header, including the private tags a scanner vendor added for its own use.
Change every named attribute
Every attribute the profile names gets removed or replaced, dates get shifted, & the patient link becomes a stable pseudonym.
Load the volume once
The pixel volume loads once, & the segmentation model runs over it with the header work already done.
Write the masks as DICOM-SEG
Organ masks are written as DICOM-SEG, referencing the de-identified series they belong to.
Send study & masks together
Study & masks move to the destination together, & anything that fails a check is held back instead.
What broke on real studies & our fixes
Four things broke when real studies hit the de-identification pass, roughly in this order. Each fix reads as obvious in hindsight, though none of them was obvious at the time.
Hidden identifiers
What broke The first clean run still carried a patient name in a free-text field somebody typed into years earlier, where no tag list looks.
The fix We now read free-text fields as text instead of trusting their tag, & anything the scan doesn’t recognize holds the study.
Broken links between studies
What broke Stripping identifiers cut the link between two scans of one patient, & that pair lost most of its research value.
The fix Dates shift by a fixed offset per patient & identifiers become pseudonyms, so studies from one person still line up.
Scanner drift
What broke A model that outlined organs correctly on one machine drifted on another, because orientation & slice thickness differed.
The fix We normalize spacing & orientation before the model sees the volume, & we hold acquisitions outside what we validated.
Silent passes
What broke Studies that failed a check had nowhere to go, so the first version passed them quietly, which a privacy pipeline can never do.
The fix A held study now raises a named reason & waits for a person, & we count a quiet pass as a defect.
A 2014 radiology review found protected information in places other than the header, so a pipeline either finds it or drops the study.[3]
Date shifting isn’t a medical idea at all, because statisticians were doing it to survey records long before imaging borrowed the trick. These are also the weeks when clients ask about bringing in specialist engineers for medical imaging. An AI proof of concept on real studies earns its keep here, because it surfaces all of this before anything gets promised.
What changed after go-live?
| Dimension | Before | Now |
|---|---|---|
| Who removes identifiers | A person, study by study | The pipeline, on every study |
| When segmentation starts | After sign-off on the scrubbing | In the same pass |
| Organ masks | Drawn by a specialist | Generated, then reviewed |
| Mask format | Whatever the tool exported | Standard DICOM-SEG |
| A field nobody recognizes | Ships with the study | Holds the study |
After go-live, the Brainy Neurals pipeline runs in production inside the client’s own environment, on the medical imaging studies its scanners produce daily. Each study comes out with no patient identifiers, after a single pass over the volume, with all of its organ masks in the standard DICOM-SEG format. Nothing crosses the network edge, which was the condition the build started from.
We haven’t published a time per study or a throughput number, & there’s no accuracy figure for the computer vision model either. The client reports movement in the right direction on all three, & the pipeline logs every held study with the reason it was held.
Day to day, a study that used to wait for a person now waits for nothing. The review that remains is a clinician checking masks, which is work worth a clinician’s time.
Handovers once scheduled around somebody’s availability now run when the studies are ready. Since handover, the client has pointed the same pass at more study types, & the rules travelled with it.
Is your imaging archive ready to share safely?
Tell us how studies leave your department today, or check whether your archive is consistent enough for a model. Either route starts from the scans you already hold.
What would we do differently?
We’d change four habits from this build, dropping one move each time & repeating another. A pipeline that’s right on one scanner & wrong on another isn’t finished yet.
Start from the published profile
Avoid A hand-written tag list, even though writing one looks quicker.
Repeat Starting from the confidentiality profile the DICOM standard publishes, which already names what carries identity.
Validate every scanner before promising
Avoid Committing to a schedule after testing one machine, as we did before the older scanner’s studies arrived two weeks later.
Repeat Running real studies from every scanner before naming a date.
Decide what a held study means
Avoid Leaving the reviewer & the response time undecided, which pushed that conversation into testing instead of scoping.
Repeat Agreeing who reviews a held study, & how fast, while scoping.
Write masks viewers already open
Avoid An output that needs a conversion step, because our first masks went unopened until we changed the export.
Repeat Writing DICOM-SEG from the very first run.
Where else does this pattern fit?
Single-pass de-identification fits wherever data must be safe to share & ready to use, because it strips identity & labels what’s inside in one automated read. Porting it means a new attribute profile & a retrained model, tested on real files.
Clinical research networks
Sites send scans to a sponsor who must never see a patient. The build adds per-site pseudonyms & a fuller audit trail.
Pharmaceutical trials
Outside reviewers read scans under blinding rules. Blinding logic sits on top of the same confidentiality profile.
Insurance claim files
Claim files carry names & policy numbers in forms & free text, with dates alongside, instead of a header. The pass reads fields rather than tags, the way our AI in banking & finance work already does.
Industrial CT in manufacturing
Industrial CT scans of parts carry a customer’s identity. Defect classes replace organ classes, & the rest of the pass matches our AI in manufacturing builds.
Veterinary imaging
Owner details ride along inside animal imaging studies. Fewer scanner types & a smaller model make this the lightest port.
Questions buyers ask before they start
What does DICOM de-identification actually remove?
The process removes or replaces every attribute that points back to a patient, such as names, identifiers, dates & free-text fields. The DICOM standard publishes a confidentiality profile naming those attributes, & a pipeline works from that list.
Is a de-identified study still protected health information?
Under HIPAA, a study that meets the de-identification standard is no longer protected health information. Getting there means removing the listed identifiers, or having an expert find that the re-identification risk is very small.
Why export organ masks as DICOM-SEG?
DICOM-SEG stores each mask as a DICOM object that references the study it came from. Any medical imaging viewer or archive that speaks DICOM opens it without a conversion step, & the mask can’t drift away from its study.
How long does automating de-identification take?
A working de-identification pipeline on your own studies usually takes a few weeks. The profile, the identifier mapping & the handling of held studies each need a pass. Segmentation quality then depends on how many scanners feed the computer vision model.
How much does an automated imaging pipeline cost?
The cost of an automated imaging pipeline follows the number of scanners, the organs you need outlined & what already sits in the archive. Brainy Neurals scopes it from a short call & a look at real studies, then quotes a fixed price. An AI readiness assessment tells you first whether your archive is consistent enough.
Who is responsible for de-identification in a clinical trial?
The site sending the images is normally responsible for removing identity before anything reaches a sponsor or a platform. That’s why automated scrubbing tends to sit at the point of export, since downstream a mistake has already travelled.
What has to leave your imaging department?
Describe the handover, the scanners feeding it & where studies need to go. We read every message & reply with the next step for your archive.
Services behind this case study
Five Brainy Neurals services sat behind this imaging build, from the header rules to the team that shipped them.
Document AI services
Identifier extraction & redaction across DICOM headers, forms & free-text fields.
Computer vision development
Segmentation models trained on volumetric medical data & validated across every scanner.
AI consulting & technology selection
Architecture & stack decisions for imaging workflows that can’t leave the network.
AI proof of concept
A few weeks on your own studies, to find what the archive is actually carrying.
Hire AI developers
Medical imaging engineers who extend your team while the format fights back.
Our engagement models cover the shapes this work usually takes. The AI in healthcare page & the industries hub show where this pattern already runs, with the wider AI development services behind all of it.
Similar case studies
Another Brainy Neurals healthcare build runs clinical output through review gates before anyone relies on it.
AI Diet Assistant for Gastroenterology
Clinical dietary guidance generated under review gates, live in a healthcare setting.
Cite this case study
Mori, Prasiddh & Patel, Mitesh. AI DICOM De-Identification & Organ Masks in Medical Imaging. Brainy Neurals, September 2026. https://brainyneurals.com/case-studies/dicom-de-identification-segmentation/
Sources cited on this page
- Aryanto KYE, Oudkerk M, van Ooijen PMA. Free DICOM de-identification tools in clinical research: functioning and safety of patient privacy. European Radiology. 2015. 25(12):3685-3695. DOI 10.1007/s00330-015-3794-0.
- Moore SM, Maffitt DR, Smith KE, Kirby JS, Clark KW, Freymann JB, Vendt BA, Tarbox LR, Prior FW. De-identification of Medical Images with Retention of Scientific Research Value. RadioGraphics. 2015. 35(3):727-735. DOI 10.1148/rg.2015140244. PMID 25969931.
- Robinson JD. Beyond the DICOM Header: Additional Issues in Deidentification. American Journal of Roentgenology. 2014. 203(6):W658-W664. DOI 10.2214/AJR.13.11789. PMID 25415732.








