AI DICOM De-Identification & Organ Masks in Medical Imaging

Home / Case studies / AI DICOM De-Identification & Organ Masks in Medical Imaging

Case study · Healthcare · Computer vision & document AI

AI DICOM De-Identification & Organ Masks in Medical Imaging

A US medical imaging organization needed DICOM de-identification on every CT & MRI study it shares with researchers, & staff handled each one manually. Brainy Neurals built a one-pass pipeline that uses a computer vision model to outline organs while it strips patient identifiers. The pipeline runs inside the client’s own network, reads each study once & replaced the manual header check before every handover. Each study now leaves with its patient identifiers removed & its organ masks attached as a standard DICOM-SEG file.

  • #MedicalImaging
  • #ComputerVision
  • #DocumentAI
  • #DICOMDeIdentification
  • #OrganSegmentation

Inside the network

Studies never leave site

Same read

Scrub & segment together

DICOM-SEG

Masks open in current viewers

Mitesh

Published October 2026

At a glance

What problem did this solve?

A US medical imaging organization removed patient identifiers from CT & MRI studies by hand. That scrubbing held up every research handover, & organ segmentation then started from scratch.

What did Brainy Neurals build?

Brainy Neurals built an automated DICOM de-identification & segmentation pipeline that runs inside the client’s own network. One pass scrubs the header against the DICOM confidentiality profile, & a computer vision model writes the organ masks as DICOM-SEG.

What changed after it went live?

Studies now leave the department with identifiers stripped & organ masks attached. Nobody retypes a header, & segmentation no longer waits for a specialist.

Who else could use this?

Any group that moves imaging studies between systems, sites or partners can use single-pass de-identification. Hospitals, research networks, imaging vendors & trial sponsors all fit.

Engagement facts

Fact Detail
Industry Healthcare, medical imaging
Client type US medical imaging organization
Engagement Imaging compliance platform
Timeline Not disclosed
Capabilities Computer vision, document AI
Delivery model Project-based delivery

Why does DICOM de-identification go wrong?

DICOM de-identification goes wrong when people do it by hand on top of a full clinical workload. A US medical imaging organization sends CT & MRI studies to researchers, vendors & partner sites, & every study has to lose its patient identity before it travels.

Technologist's hands transcribing a printed study worksheet beside a stack of archived imaging discs
Before the pipeline, a person read every study line by line & struck out anything that named the patient.

Scrubbing a study is document AI work, & most of it happens in the header. A DICOM header is the block of labelled fields, called tags, stored with every scan. Headers run long & drift into free text, & each scanner writes them its own way.

Field by field

A technician read each header & typed over any name they found.

One miss, one rework

A single name left in a description field sent a whole handover back for review.

Two queues

Segmentation waited until the scrubbing was signed off, so the two jobs never overlapped.

Slice by slice

A specialist drew organ outlines by hand while real cases waited in the queue.

Every scanner differs

Each machine wrote its header its own way, so last month’s rule failed this week.

Ten toolkits, one safe default

A 2015 evaluation tested 10 free DICOM de-identification toolkits for this weakness. Only one removed every required element on its default settings.[1]

Why do common approaches stall?

Common approaches stall because each of the four routes to scrubbing a study only fits some settings. By hand, tiredness decides what gets missed, & the other three stop at a wall the fourth one doesn’t meet.

Approach Strength Where it stops Still suits
Scrub it by hand Human judgment on every field One study at a time Rare, small handovers
Off-the-shelf anonymizer Covers the obvious tags fast Misses private tags & free text Single-scanner archives
Delete everything optional Nothing identifying survives export Dates & links break research value One-off dataset drops
One automated pass (our route) Identity & masks handled together Weeks of profile & validation work Studies moving in volume

A 2015 review notes that deleting all private elements removes useful scientific data along with the identifiers.[2]

What Brainy Neurals built

Brainy Neurals built a one-pass imaging pipeline for a US medical imaging organization, so each study is read once & written once. The header work & the organ segmentation run over the same loaded volume on the client’s own network. Scrubbing & segmentation stopped being two queues waiting on each other.

Architecture diagram of the Brainy Neurals pipeline. Inside the client network, a watched intake passes each study to a header read that applies rules from the DICOM attribute profile and writes pseudonyms to an identifier map. The pixel volume then loads once, an organ model segments it, the mask is written as DICOM-SEG, and study plus mask travel to the destination together. A study with an unknown field or an out-of-range acquisition is held. Nothing outside the network is in the loop. THE CLIENT NETWORK DE-IDENTIFY SEGMENT 10 Outside network NOT IN THE LOOP 1 Watched intake 3 Attribute profile RULES 2 Header read PSEUDONYMS 4 Identifier map 5 Volume loaded 6 Organ model 7 Mask written DICOM-SEG 8 Destination UNKNOWN FIELD OUT OF RANGE 9 Held studies Architecture diagram of the Brainy Neurals pipeline. Inside the client network, a watched intake passes each study to a header read that applies rules from the DICOM attribute profile and writes pseudonyms to an identifier map. The pixel volume then loads once, an organ model segments it, the mask is written as DICOM-SEG, and study plus mask travel to the destination together. A study with an unknown field or an out-of-range acquisition is held. Nothing outside the network is in the loop. THE CLIENT NETWORK 1 Watched intake 2 Header read 3 Attribute profile RULES 4 Identifier map PSEUDONYMS 5 Volume loaded 9 Held studies UNKNOWN FIELD OR OUT OF RANGE 6 Organ model 7 Mask written DICOM-SEG 8 Destination 10 Outside network NOT IN THE LOOP

One pass over one loaded volume, & nothing crosses the client’s network edge.

The first decision was where the rules come from, & we started from the confidentiality profile the DICOM standard publishes. A hand-kept tag list only knows the tags it has already met. We map identifiers & dates rather than deleting them. A study that loses its dates loses the reason somebody wanted it, so that choice was deliberate.

A study that loses its dates loses the reason somebody wanted it, so that choice was deliberate.

The second decision was what comes out the other end. Masks are written as DICOM-SEG against the de-identified study, so they open in the viewers the client already runs. We turned down a research-format side file, because a side file gets separated from its study within a week.

Teaching a model to outline organs is ordinary computer vision development, & everything around the model is what makes its output usable.

The stack we used

The stack runs inside the client’s own environment, because none of this data may leave it. That ruled out a hosted processing service before we chose anything else. We picked each remaining layer for what it costs to run & maintain, which is technology selection work done before any code.

DICOM read & written natively

The study keeps its structure from intake to export. We ruled out converting to a research format.

The standard’s confidentiality profile

The DICOM standard already names what carries identity. We ruled out a hand-kept deny-list.

Pseudonyms & shifted dates

Studies from one patient still group after the names go. We ruled out fresh random values.

Medical imaging deep-learning framework

It’s built for volumetric data, which is what CT & MRI scans produce. We ruled out a general vision library.

DICOM-SEG against the source series

Masks open in the viewers the client already runs. We ruled out a mask file sitting beside the study.

Python services with a watched intake

Python keeps one language across the whole toolchain. We ruled out a workflow product that needs a license.

How does one study move through the pipeline?

One study moves through six steps in the order the pipeline sees it, & a person appears in none of them.

Check the study at intake

A study lands in the watched intake, & the pipeline checks it’s complete before touching anything.

Read the full header

The pipeline reads the full header, including the private tags a scanner vendor added for its own use.

Change every named attribute

Every attribute the profile names gets removed or replaced, dates get shifted, & the patient link becomes a stable pseudonym.

Load the volume once

The pixel volume loads once, & the segmentation model runs over it with the header work already done.

Write the masks as DICOM-SEG

Organ masks are written as DICOM-SEG, referencing the de-identified series they belong to.

Send study & masks together

Study & masks move to the destination together, & anything that fails a check is held back instead.

What broke on real studies & our fixes

Four things broke when real studies hit the de-identification pass, roughly in this order. Each fix reads as obvious in hindsight, though none of them was obvious at the time.

Archive shelf of mismatched imaging media and a legacy scanner console panel in a storage room
Studies arrive from scanners of different ages, & no two of them write a header the same way.

Hidden identifiers

What broke The first clean run still carried a patient name in a free-text field somebody typed into years earlier, where no tag list looks.

The fix We now read free-text fields as text instead of trusting their tag, & anything the scan doesn’t recognize holds the study.

Broken links between studies

What broke Stripping identifiers cut the link between two scans of one patient, & that pair lost most of its research value.

The fix Dates shift by a fixed offset per patient & identifiers become pseudonyms, so studies from one person still line up.

Scanner drift

What broke A model that outlined organs correctly on one machine drifted on another, because orientation & slice thickness differed.

The fix We normalize spacing & orientation before the model sees the volume, & we hold acquisitions outside what we validated.

Silent passes

What broke Studies that failed a check had nowhere to go, so the first version passed them quietly, which a privacy pipeline can never do.

The fix A held study now raises a named reason & waits for a person, & we count a quiet pass as a defect.

A 2014 radiology review found protected information in places other than the header, so a pipeline either finds it or drops the study.[3]

Date shifting isn’t a medical idea at all, because statisticians were doing it to survey records long before imaging borrowed the trick. These are also the weeks when clients ask about bringing in specialist engineers for medical imaging. An AI proof of concept on real studies earns its keep here, because it surfaces all of this before anything gets promised.

What changed after go-live?

Dimension Before Now
Who removes identifiers A person, study by study The pipeline, on every study
When segmentation starts After sign-off on the scrubbing In the same pass
Organ masks Drawn by a specialist Generated, then reviewed
Mask format Whatever the tool exported Standard DICOM-SEG
A field nobody recognizes Ships with the study Holds the study
Technologist preparing a scanner table with no paperwork in reach after imaging workflow automation
The same desk after go-live, with each study loaded once & the retyping gone from the job.

After go-live, the Brainy Neurals pipeline runs in production inside the client’s own environment, on the medical imaging studies its scanners produce daily. Each study comes out with no patient identifiers, after a single pass over the volume, with all of its organ masks in the standard DICOM-SEG format. Nothing crosses the network edge, which was the condition the build started from.

We haven’t published a time per study or a throughput number, & there’s no accuracy figure for the computer vision model either. The client reports movement in the right direction on all three, & the pipeline logs every held study with the reason it was held.

Day to day, a study that used to wait for a person now waits for nothing. The review that remains is a clinician checking masks, which is work worth a clinician’s time.

Handovers once scheduled around somebody’s availability now run when the studies are ready. Since handover, the client has pointed the same pass at more study types, & the rules travelled with it.

Is your imaging archive ready to share safely?

Tell us how studies leave your department today, or check whether your archive is consistent enough for a model. Either route starts from the scans you already hold.

What would we do differently?

We’d change four habits from this build, dropping one move each time & repeating another. A pipeline that’s right on one scanner & wrong on another isn’t finished yet.

Start from the published profile

Avoid A hand-written tag list, even though writing one looks quicker.

Repeat Starting from the confidentiality profile the DICOM standard publishes, which already names what carries identity.

Validate every scanner before promising

Avoid Committing to a schedule after testing one machine, as we did before the older scanner’s studies arrived two weeks later.

Repeat Running real studies from every scanner before naming a date.

Decide what a held study means

Avoid Leaving the reviewer & the response time undecided, which pushed that conversation into testing instead of scoping.

Repeat Agreeing who reviews a held study, & how fast, while scoping.

Write masks viewers already open

Avoid An output that needs a conversion step, because our first masks went unopened until we changed the export.

Repeat Writing DICOM-SEG from the very first run.

Where else does this pattern fit?

Single-pass de-identification fits wherever data must be safe to share & ready to use, because it strips identity & labels what’s inside in one automated read. Porting it means a new attribute profile & a retrained model, tested on real files.

Clinical research networks

Sites send scans to a sponsor who must never see a patient. The build adds per-site pseudonyms & a fuller audit trail.

Pharmaceutical trials

Outside reviewers read scans under blinding rules. Blinding logic sits on top of the same confidentiality profile.

Insurance claim files

Claim files carry names & policy numbers in forms & free text, with dates alongside, instead of a header. The pass reads fields rather than tags, the way our AI in banking & finance work already does.

Industrial CT in manufacturing

Industrial CT scans of parts carry a customer’s identity. Defect classes replace organ classes, & the rest of the pass matches our AI in manufacturing builds.

Veterinary imaging

Owner details ride along inside animal imaging studies. Fewer scanner types & a smaller model make this the lightest port.

Questions buyers ask before they start

What does DICOM de-identification actually remove?

The process removes or replaces every attribute that points back to a patient, such as names, identifiers, dates & free-text fields. The DICOM standard publishes a confidentiality profile naming those attributes, & a pipeline works from that list.

Is a de-identified study still protected health information?

Under HIPAA, a study that meets the de-identification standard is no longer protected health information. Getting there means removing the listed identifiers, or having an expert find that the re-identification risk is very small.

Why export organ masks as DICOM-SEG?

DICOM-SEG stores each mask as a DICOM object that references the study it came from. Any medical imaging viewer or archive that speaks DICOM opens it without a conversion step, & the mask can’t drift away from its study.

How long does automating de-identification take?

A working de-identification pipeline on your own studies usually takes a few weeks. The profile, the identifier mapping & the handling of held studies each need a pass. Segmentation quality then depends on how many scanners feed the computer vision model.

How much does an automated imaging pipeline cost?

The cost of an automated imaging pipeline follows the number of scanners, the organs you need outlined & what already sits in the archive. Brainy Neurals scopes it from a short call & a look at real studies, then quotes a fixed price. An AI readiness assessment tells you first whether your archive is consistent enough.

Who is responsible for de-identification in a clinical trial?

The site sending the images is normally responsible for removing identity before anything reaches a sponsor or a platform. That’s why automated scrubbing tends to sit at the point of export, since downstream a mistake has already travelled.

What has to leave your imaging department?

Describe the handover, the scanners feeding it & where studies need to go. We read every message & reply with the next step for your archive.







    Services behind this case study

    Five Brainy Neurals services sat behind this imaging build, from the header rules to the team that shipped them.

    Document AI services

    Identifier extraction & redaction across DICOM headers, forms & free-text fields.

    Computer vision development

    Segmentation models trained on volumetric medical data & validated across every scanner.

    AI consulting & technology selection

    Architecture & stack decisions for imaging workflows that can’t leave the network.

    AI proof of concept

    A few weeks on your own studies, to find what the archive is actually carrying.

    Hire AI developers

    Medical imaging engineers who extend your team while the format fights back.

    Our engagement models cover the shapes this work usually takes. The AI in healthcare page & the industries hub show where this pattern already runs, with the wider AI development services behind all of it.

    Similar case studies

    Another Brainy Neurals healthcare build runs clinical output through review gates before anyone relies on it.

    AI Diet Assistant for Gastroenterology

    Clinical dietary guidance generated under review gates, live in a healthcare setting.

    Cite this case study

    Mori, Prasiddh & Patel, Mitesh. AI DICOM De-Identification & Organ Masks in Medical Imaging. Brainy Neurals, September 2026. https://brainyneurals.com/case-studies/dicom-de-identification-segmentation/

    Sources cited on this page

    1. Aryanto KYE, Oudkerk M, van Ooijen PMA. Free DICOM de-identification tools in clinical research: functioning and safety of patient privacy. European Radiology. 2015. 25(12):3685-3695. DOI 10.1007/s00330-015-3794-0.
    2. Moore SM, Maffitt DR, Smith KE, Kirby JS, Clark KW, Freymann JB, Vendt BA, Tarbox LR, Prior FW. De-identification of Medical Images with Retention of Scientific Research Value. RadioGraphics. 2015. 35(3):727-735. DOI 10.1148/rg.2015140244. PMID 25969931.
    3. Robinson JD. Beyond the DICOM Header: Additional Issues in Deidentification. American Journal of Roentgenology. 2014. 203(6):W658-W664. DOI 10.2214/AJR.13.11789. PMID 25415732.