An AI model can produce a convincing demonstration: an image arrives, the software highlights a region and a result appears. Turning that demonstration into a useful medical imaging product requires a much broader set of decisions.
Who will use the result? Which patients, scanners and image types does the evidence cover? What happens when the software cannot process a study, or when its output is wrong?
AI in medical imaging can support tasks such as identifying patterns, outlining structures and assisting imaging workflows. Its value depends on the specific task, the quality of the evidence and how the system fits into clinical practice.
For founders and product teams, the starting point is a clearly defined purpose and a plan to demonstrate that the product can fulfil it. This is a product-development overview, not clinical advice or a complete regulatory guide. It does not claim that Addbox has delivered regulated imaging devices. Examples and numbers are illustrative.
Choose a specific problem to solve
“AI for radiology” is too broad to guide development. An application designed to outline a structure for review has different requirements from one designed to prioritise studies or assist interpretation.
| Potential application | Intended contribution | Key question to investigate |
|---|---|---|
| Image quality assessment | Flag images that may be unsuitable for a particular analysis | Does the flag identify relevant quality problems reliably? |
| Segmentation | Outline a defined structure for a qualified user to inspect | How much correction is needed, and what are the consequences of an error? |
| Measurement assistance | Produce a defined measurement from an image | Is it reproducible and sufficiently accurate for its intended use? |
| Detection support | Draw attention to a specified finding | What is missed, what is incorrectly flagged and how does this affect review? |
| Worklist prioritisation | Help determine the order in which studies are reviewed | What happens to both flagged and unflagged cases? |
| Workflow administration | Route studies or identify missing information | Does it reduce delays without creating mismatches or lost work? |
These are categories of possible applications, not claims that a particular model is suitable for them. Each needs an appropriate evaluation.
Work with clinicians who understand the proposed setting. A technically interesting task may contribute little if it addresses a step that is already quick or moves the bottleneck elsewhere.
Write the intended purpose before choosing the model
Describe the input, output, intended users, patient population and setting. Explain how the result should influence the next action, including what the product is not intended to do.
The MHRA's guidance on defining intended purpose for software as a medical device connects this definition to the evidence needed to support the product. Broad or vague claims make meaningful evaluation difficult.
An initial brief might specify one imaging modality, a defined examination type and a particular professional user. The team can then investigate whether the proposed data and studies support that scope.
Adding more populations, image types or clinical uses later is a change to the product's claims. It should trigger a review of the evidence and associated risks, rather than being treated as an ordinary feature expansion.
Distinguish a prototype from a clinical product
A research prototype may answer whether an approach is technically promising. It does not establish that the software is ready to influence care.
| Stage | Question being answered | What the result does not establish on its own |
|---|---|---|
| Technical experiment | Can the proposed method process the input and produce the required output? | Clinical validity or suitability for routine use. |
| Retrospective evaluation | How does the fixed system perform on appropriately selected existing data? | Performance in the full live workflow. |
| Prospective evaluation | How does it behave under a planned real-world study? | Permission for every intended deployment or claim. |
| Operational deployment | Can the authorised product be operated reliably in its intended setting? | That future updates will preserve its performance. |
The appropriate sequence, approvals and evidence depend on the intended use and jurisdiction. Calling a product a prototype, or adding a clinician review step, does not automatically remove medical-device obligations.
Use the MHRA's software and AI medical-device resources to orient the project, and establish the applicable route with qualified regulatory and clinical specialists before clinical use.
Build a data plan that matches the intended setting
The number of images is only one consideration. Ask where the data came from, which people it represents and whether the acquisition conditions resemble the proposed deployment.
Relevant differences may include scanners, sites, imaging protocols, patient characteristics and the frequency of the finding being assessed. A dataset selected because it is convenient may not answer the question your product needs to answer.
Define how reference labels or measurements are established. If professionals disagree, specify how those disagreements are resolved and preserve uncertainty where appropriate.
Keep training and tuning separate from final testing. The FDA's digital health glossary explicitly distinguishes independent test data from data used during development.
For imaging projects, check for repeated patients, related examinations and other overlap that could make a test easier than a genuinely independent evaluation. Agree the evaluation design with clinical and statistical expertise before repeatedly inspecting results.
Document permissions and permitted uses for the data. De-identification also needs a considered process: identifying information can occur in metadata and image content, so removing a visible name is not a complete method.
Evaluate errors in context
A single headline accuracy figure can conceal the errors that matter most.
For a detection task, sensitivity describes how often relevant positive cases are identified, while specificity describes how often negative cases are correctly identified. The usefulness of a positive flag also depends on how common the finding is in the evaluated population.
Consider a purely illustrative set of 1,000 examinations, with 100 positive and 900 negative cases. At 90% sensitivity and 90% specificity, the system would identify 90 positive cases, miss 10 and incorrectly flag 90 negative cases. Only half of its 180 positive flags would be true positives.
This arithmetic is not a proposed clinical target. It illustrates why a percentage without context is insufficient for assessing a product.
Predefine task-appropriate measures, quantify uncertainty and examine performance across relevant groups and settings. For segmentation, review whether errors affect the intended task as well as numerical overlap. For prioritisation, examine delays experienced by cases the system does not flag.
Evaluate the human and software together where the intended use involves professional review. A technically strong model may still create unnecessary work or encourage users to overlook information outside its output.
Design the integration as part of the product
The software needs to receive the correct study and return a result that remains associated with the correct patient, examination and version.
DICOM is the standard for communicating medical images and associated information. Supporting a standard is a starting point for interoperability; actual integration still needs testing against the systems and configurations in use.
An imaging product may need to work with a picture archiving and communication system, or PACS, as well as other clinical systems. Determine where processing occurs, how results are presented and who is notified when processing fails.
The design should account for incomplete transfers, duplicate submissions, corrected information, unsupported inputs and downtime. “No result produced” must remain distinguishable from a valid result indicating no finding.
Clinical users need an appropriate fallback. A service outage should not silently remove an examination from the normal review process.
Make the interface support appropriate review
Show what the output relates to, whether processing completed and any relevant limitations. Let users inspect the underlying image and correct or reject outputs where the workflow requires it.
The FDA, Health Canada and MHRA's transparency principles for machine-learning-enabled devices emphasise information that helps users understand and use a device appropriately, including performance and limitations.
Do not assume that an attractive overlay explains why a result is correct. Test whether the presentation helps the intended users complete their task and recognise when the software is unsuitable.
If generated text is included, evaluate it separately. A fluent report can contain unsupported statements or omit important context. Keep measurements, verified observations and generated language distinguishable in the design.
Plan for changes after launch
A medical imaging product is an ongoing operational responsibility. Model updates, changes in input data and modifications to connected systems can affect behaviour.
The FDA's Good Machine Learning Practice page points to the IMDRF's lifecycle principles for developing AI-enabled medical devices. The lifecycle perspective matters: evidence and oversight continue after the first release.
Maintain traceable versions of the model, preprocessing, configuration and application. Define what monitoring can detect, who investigates problems and how an affected release can be withdrawn or rolled back.
Successful processing is not proof of clinical correctness. Operational monitoring and clinical performance review answer different questions, and both need an appropriate plan.
Do not introduce automatic retraining or unreviewed model substitutions into a deployed product as a routine optimisation. Assess proposed changes against the validation and regulatory arrangements for that product.
Budget for the work beyond the model
An early project plan should identify the resources required for data access, annotation, clinical participation, evaluation, integration, quality management and support.
Useful questions for a development brief:
- What exact claim will the first version make?
- Who provides clinical, statistical and regulatory expertise?
- Which data can be used, and for which purposes?
- What evidence is needed before the next stage?
- Which systems and sites must the product work with?
- Who monitors the deployed service and handles incidents?
- What happens if the evidence does not support the intended claim?
A smaller, well-defined purpose gives the team a clearer basis for estimating this work. Building a broad demo first can defer the most consequential questions until substantial effort has already been spent.
Start with the purpose and the evidence
AI in medical imaging offers opportunities where a specific task can be supported by reliable evidence and a workable clinical process. The development challenge includes the model, the interface, the connected systems and the arrangements for operating the product over time.
For teams exploring the software architecture and integration work around an imaging project, see Addbox's development services or AI strategy service. Establish the clinical and regulatory workstreams alongside the technical scope from the outset.