WritingResearch systems

From a device export to a research system

How I built and operate CorneaForge / OphtaFlow AI: device ingestion, scientific processing, examination review, and model serving on hospital infrastructure.

In this article

I built CorneaForge to connect ophthalmic devices, scientific processing, and research in one on-premise system at Quinze-Vingts Hospital. Its application, OphtaFlow AI, lets users inspect examinations, measurements, images, native videos, annotations, and existing model predictions.

That work includes operating the system. Device exports have to arrive, background computations have to finish, and an examination opened in the interface has to connect to the right source data. The research experiments depend on those same foundations.

A read-only inspection on 1 October 2026 confirmed the running deployment and its outputs: MS-39 indexing, feature extraction and rendering; OCT processing; Corvis DICOM reception and promotion into structured records; the application API; PostgreSQL; and MinIO. It also verified existing model-result serving and a separate working XGBoost inference route. This article follows that deployed architecture and the decisions behind it.

Separate storage, processing, and interactive work

The core stack is straightforward. MinIO stores objects such as original device exports and derived files. PostgreSQL holds structured measurements, acquisition relationships, annotations, and processing state. FastAPI exposes the application interfaces. Dedicated workers perform indexing, feature computation, rendering, and OCT processing; systemd supervises them.

Device exportsMS-39 and Corvis
Retained sourcesMinIO objects and native indexes
Background processingFeatures, maps, and OCT sections
Structured acquisitionsPostgreSQL measurements and state
OphtaFlow AIReview, annotation, model results
Research datasetsStudy selection, Parquet, DuckDB
Schematic of the deployed data path and its application and research consumers. Systemd supervises the API and processing workers.

I separated these responsibilities because they have different execution patterns. Opening an examination needs a responsive request path. Decoding an export, generating maps, or processing an OCT series may take longer and should be recoverable independently. A worker can keep processing while the application serves already available results.

The deployment uses systemd resource groups and CPU allocations to separate interactive application work from heavier research processing. For example, the API and integrated AutoMorph worker occupy different CPU sets. This is a practical way to operate the platform alongside computational research on shared hospital hardware.

Retain the source, then build the useful views

A device export is the beginning of a lineage: the chain connecting a derived result to its inputs and transformations. Retaining the original objects lets me inspect a parser disagreement, improve a computation, or rebuild a derived view without asking the instrument to reproduce an old export.

CorneaForge organizes that work into three layers:

LayerResponsibilityConcrete implementation
BronzePreserve and index received source dataDevice objects in MinIO and native source indexes
SilverStructure acquisitions and their measurementsPostgreSQL acquisition records, device features, annotations, and processing state
GoldAssemble data for a defined studyResearch tables written as Parquet and queried with DuckDB

The acquisition is the link between these layers. Device-specific identifiers and source references let the system associate files and derived results with an examination, while preserving the device and laterality needed to interpret it.

Keeping the received object also makes retransmission manageable. A repeated export should resolve through its source identity rather than create a new examination simply because a file arrived twice. That principle matters for ordinary network retries as much as for historical imports.

Bring device data into the application

I built separate background services to collect MS-39 data, extract measurements from device files, and generate corneal maps. The application can then display these results without repeating the calculations every time an examination is opened.

For OCT—the cross-sectional scans of the front of the eye—the system records where each image is stored and how it should be displayed. The viewer renders images on demand and caches them for subsequent views.

The Corvis pipeline receives device exports through DICOM and imports their measurements into OphtaFlow AI. Incoming files and the imported results remain linked, so I can trace a displayed measurement back to its source.

Preserve what each measurement means

Ophthalmic processing connects several kinds of numbers: native device measurements, quantities recomputed from surfaces, and values used to render an image. I keep their origins identifiable because they answer different questions.

For example, a native corneal index can depend on the device’s own analysis. A geometric quantity computed from exported points depends on our fitting and sampling conventions. A colored map depends on interpolation and a display scale. A similar label does not make these outputs equivalent.

In the geometry code, this led me to separate reusable spatial structure from changing measured values. For repeated polar-to-Cartesian maps, I cache triangle lookups and interpolation weights when their geometry is compatible. Missing source points participate in the cache identity because they can change that geometry. I explain the implementation in Cache the geometry, then change the values, and the surface-fitting conventions in From measured points to Zernike coefficients.

Processing state carries meaning too. A failed source read, a confirmed absent object, and a parsing failure need different responses. The MS-39 feature path distinguishes these outcomes and preserves valid earlier features when a later extraction attempt fails. This makes retry behavior useful without turning temporary storage trouble into an apparently empty examination.

Serve the models that already exist

The deployed application returns stored predictions from six model heads, covering a legacy multiclass model and five binary axes. During inspection, an acquisition request successfully returned those results alongside its features. The database contained 62,635 acquisitions per head. A separate deployed seven-class XGBoost demo route performed inference successfully as well.

There is a specific freshness boundary: the latest stored predictions are dated 8 August 2026, and the background MS-39 prediction worker is currently disabled. The application serves those existing results; this does not imply that every newly ingested examination receives a fresh prediction. Neither the new governed activation workflow nor future JEPA integration is needed to explain the serving capabilities already present.

The engineering requirement for the next model remains concrete: deploy the preprocessing, ordered feature schema, units, and weights as a compatible package. Changing the meaning of an input while retaining its column name can break a model just as surely as changing the number of columns. The AS-OCT encoders I am training are research candidates for a later integration, with their own evaluation work. Serving a model and establishing its clinical performance are separate achievements.

Turn operational data into an experiment

A working database is useful for browsing examinations. A research dataset additionally defines a population, a row unit, selected variables, labels, and a construction version. CorneaForge’s research builders produce Parquet tables and use DuckDB for analysis, keeping analytical scans away from the transactional database. A read-only query confirmed an existing snapshot containing MS-39, Corvis, and imported IOLMaster biometry. That snapshot was built on 25 June; it remains separate from the live ingestion state.

The row unit matters when joining devices. A table of acquisitions is different from a table of eyes or people. Repeated examinations can otherwise multiply rows and change the weight of a person in an analysis. Dataset construction is where I make that choice explicit and retain the references needed to recover its source material.

This infrastructure supports my AS-OCT pretraining experiments, which use a fixed quality-filtered index and recorded source order. The resulting representation measurements belong to a dated research snapshot; new hospital acquisitions do not silently rewrite that comparison.

I also integrated AutoMorph, the retinal analysis pipeline, as a separate upload, queue, and download service. Its interface and worker were running during inspection, with the worker idle. Separately, the completed CLARUS batch research run has a verified output manifest spanning 154 batches. The retained web queue does not establish a recent completed submission, so those two operational records remain distinct.

What operating the system changes

Building the models and operating their data system makes the engineering questions specific. I need to know which source produced a surprising value, which transformation to repeat, whether a worker can recover, and which dataset an experiment actually used.

CorneaForge gives those questions concrete objects: retained exports, structured acquisitions, independently supervised workers, inspectable application results, and versioned research outputs. That is the connection I wanted between hospital data and research: a system I can use, investigate, and extend while the experiments evolve.

The deployment measurement record contains the dated aggregate counts and verification scope used here. It includes no examination records, images, credentials, or internal network addresses.