In this article
An OCT cross-section can contain a curved band of anatomy surrounded by a large amount of background. Before choosing a model, there is already an engineering question: where should the computation go?
One approach reconstructs missing parts of the whole image. Another spends a limited patch budget around a segmented anatomical region. A third learns to remove noise. A fourth asks global and local views to agree in representation space.
I implemented five AS-OCT recipes to test these choices with a shared image index and encoder initialization. This article records the configuration established on 20 September 2026, including the matched-patch settings reviewed on 23 September. The first measured feature spectra are in the accompanying research note.
Start by deciding what one example means
Optical coherence tomography, or OCT, produces cross-sectional images. A single cross-section is a B-scan. An examination contains several B-scans, and one person may have several examinations. Patches and augmented views are derived from those images.
I count source-image presentations for pretraining and group downstream evaluation by patient, so repeated scans and views retain their relationship to the same person.
The recorded training index contains 3,427,130 B-scans after the shared quality filter. Its label, QC085, means that the recorded bad-quality score must be strictly below 0.85. The score is an uncalibrated quality ranking; its definition and the retained distribution are documented in the research note.
All five recipes use the same source index and recorded source order. Each update represents 64 source images; the scheduled budget is 150,000 updates. Therefore:
150,000 updates × 64 sources = 9,600,000 image presentations
Presentations include repeated images. The sampling unit is the B-scan, so people with more scans contribute more presentations. That choice should remain visible when interpreting the experiment.
Across this index, 9.6 million presentations correspond to about 2.8 passes over the source images. Producing ten views of each source increases the computation; it does not turn those 2.8 passes into 28 epochs. I use this fixed source-exposure budget to compare the complete recipes; a compute-matched comparison would answer a different question.
The encoders also share a starting point: the same ImageNet-pretrained MAE ViT-L/16 checkpoint. These are experiments in continued domain pretraining. They are not five models trained from random initialization, and their decoders or projectors do not inherit an identical training history merely because the encoders do.
ViT-L/16 denotes a large vision transformer with 16-by-16-pixel patches. A patch is a computational unit, not a measure of the scanner’s optical resolution.
| Input views | Encoder input | Target |
|---|---|---|
| Full-field MAE Full image, with each spatial dimension divided by 2. | 15% of patches visible; 85% masked. | Reconstruct normalized pixel values at the masked positions. |
| Matched-Global Image dimensions divided by 2. Select up to 128 patches from global positions. | The shared budget gives at most 32 visible patches. The remaining selected patches are masked. | Reconstruct normalized pixel values at the selected target positions. |
| Matched-CUNEX Same image scale and patch budget; select positions using the anatomical mask. | The same visible and target counts as Matched-Global; only the selected positions change. | Reconstruct normalized pixel values at the selected target positions. |
| DiffMAE 2D Full image, with each spatial dimension divided by 2. | 25% clean visible patches. The decoder also receives noisy hidden content at the other 75%. | Predict the clean normalized target, x₀, at the hidden positions. |
| JEPA + VISReg Two global views with dimensions divided by 3, plus eight native 512 × 512 crops: six mask-guided and two free. | A shared encoder processes all patches of each view, followed by a shared projector. | Agreement with the mean global projection, plus VISReg regularization. |
Shared matched-patch budget: M = min(128, C), V = ceil(M / 4), T = M − V. C is the number of anatomical candidate patches; both matched arms use the same counts. M counts selected patches, and V counts the patches visible to the encoder.
Recipe 1: reconstruct the full field
The first arm is the full-field MAE baseline. The entire raster is reduced by a factor of two in each spatial dimension, with its aspect ratio preserved. Then 85% of its patches are hidden. The encoder sees the remaining patches; the decoder predicts normalized pixel targets in the hidden regions.
This adapts the masked reconstruction mechanism introduced by He and colleagues. The original paper prominently used 75% masking; the 85% setting here belongs to this experiment. “Canonical MAE,” the internal run name, therefore should not be read as “every setting exactly reproduces the paper.” MAE paper
The appeal is clear: the encoder receives distributed context across the image. The cost is also clear: the learning problem includes whatever appears in that field, including background and acquisition characteristics. Reconstruction does not specify which image details a later diagnostic task should prioritize.
The question this arm asks is practical: what representation does a full-field reconstruction recipe learn under the shared data and update budget?
Recipes 2 and 3: hold the patch budget fixed, change its placement
The next pair makes one comparison more controlled. Both use the same reduced image and the same number of selected, visible and target patches. They differ in where positions are selected.
Matched-Global draws from the image grid. Matched-CUNEX draws from candidates identified using an anatomical segmentation mask. Here, CUNEX names the segmentation used to guide selection. The mask identifies a region; it does not label a disease or guarantee that every useful finding lies inside it.
For the recorded rule, a patch becomes an anatomical candidate when more than 10% of its area overlaps the mask. Let C be the number of these candidates. The shared budget is:
M = min(128, C) selected patches
V = ceil(M / 4) visible patches
T = M − V reconstruction targets
If C = 100, both arms select 100 patches: 25 visible and 75 targets. If C = 200, both select 128: 32 visible and 96 targets. The global arm may have many more available positions, but it does not receive a larger budget because of that.
C = 100; selected M = 100; visible V = 25; target T = 75.
M = min(128, C), V = ceil(M / 4), T = M − V.
The matched recipes share this rule. Their candidate locations differ. The bar describes a token allocation, not accuracy or equal compute.
Change the candidate count to see where the budget stops growing. This is an illustration of the selection rule, using constructed numbers; it does not display images or results from the study.
Three sets are easy to confuse here: the complete image grid, the selected patches and the patches actually passed to the encoder. A drawing that colors all selected patches as visible would misrepresent the experiment.
Positions remain unique; missing capacity is not filled by duplicating patches. The recorded implementation treats fewer than two anatomical candidates as an error rather than silently removing that source from one arm’s dataset.
The hypothesis is that spatial guidance can spend computation more usefully. Its counterargument deserves equal attention: segmentation errors can systematically hide useful regions. Even a correct anatomical mask may omit context that matters. Mask-guided selection changes what is sampled; it does not replace image intensities with a silhouette.
Because these two arms share more ingredients, their comparison is more informative about spatial selection than a comparison between entirely different objectives and view geometries.
Recipe 4: reconstruct by denoising
The DiffMAE arm changes the decoder’s task. Visible regions provide clean context, while hidden regions enter the decoder in a noisy form. The decoder also receives information about the noise level and learns to recover a clean target. This follows the masked, conditional diffusion approach introduced by Wei and colleagues. DiffMAE paper
Our recorded adaptation is two-dimensional, uses 75% masking and predicts the clean normalized target, commonly called x0. It uses the same full-field reduction by two as the reconstruction arms.
The noise schedule has 1,000 levels. That does not mean the training loop performs 1,000 successive denoising operations for each update. A level is sampled to construct the training example. The schedule specifies possible corruption strengths; an iterative generation procedure is a different computation.
Imagine hiding part of our synthetic arc. Plain reconstruction supplies the missing positions and asks for their values. The denoising recipe also supplies a corrupted version of that hidden content and asks the decoder to correct it using context and the noise level.
This creates additional modeling capacity and cost. Whether the encoder learns more useful features is still a question for a common downstream evaluation. A more elaborate decoder or a better reconstructed image does not answer it on its own.
Recipe 5: combine global context with native local views
The JEPA+VISReg arm changes both the target and the way the image is presented. Its ten views are:
- Two independently augmented full-field views, each reduced by a factor of three.
- Six local windows guided by the anatomical mask.
- Two local windows selected independently of that mask.
The local windows are 512 by 512 pixels, extracted from the native corpus raster. They are not enlarged crops of a reduced global image. The independent windows provide opportunities to observe areas that mask-guided sampling might miss; they do not guarantee complete coverage.
All views pass through a shared encoder and projector. For each source, the mean of the two global projections becomes the agreement target. The objective combines 0.3 × agreement with 0.7 × VISReg, and gradients flow through the global target. There is no separate moving-average teacher in this implementation. VISReg provides explicit control over representation scale and distributional shape. VISReg paper
The intended tradeoff is to combine overall shape with local sampling detail. But a global target may fail to preserve a small local distinction. Keeping high-resolution pixels available does not force the objective to use them.
“Native” needs care too. It describes the sampling grid of the JPEG corpus, which has already been processed and recompressed. It does not mean access to raw optical measurements. Preserving a raster’s aspect ratio likewise does not make its physical sampling isotropic.
Image preparation belongs in the experiment
All four reconstruction arms preserve the full field and use only the padding needed to complete the patch grid. Artificial padding pixels are excluded from target normalization and reconstruction loss. Otherwise, changing the padding could change the objective despite leaving the observed anatomy untouched.
The recorded reconstruction losses are averaged per image and then across the 64 sources. That prevents an image with more valid target pixels from automatically receiving more weight solely because of its dimensions.
These details may look smaller than the architecture choice. In practice, they define what “the same experiment” means. A silent change to target normalization, valid-pixel handling or image weighting changes the question the model is optimizing.
Equal exposure is only one kind of fairness
The shared source order, initialization and presentation budget create useful controls. They do not equalize everything.
Ten-view JEPA and single-view reconstruction process different amounts of information. Different resolutions produce different token counts. Decoders, objectives and memory requirements differ. Equal source exposure therefore does not imply equal GPU time or equal floating-point work.
There are several legitimate experimental questions:
| Question | What needs to be controlled or reported? |
|---|---|
| What happens after the same source exposure? | Index, source order and number of presentations |
| What can each recipe achieve within a compute budget? | Hardware, elapsed time, memory and execution settings |
| Does mask-guided placement help at a fixed patch budget? | Image preparation, initialization, visible/target counts and selection rule |
A study should state which question it is answering. Comparing full-field MAE with global–local JEPA changes views, resolution and objective together. That comparison can evaluate complete recipes, but cannot isolate the causal effect of replacing one loss function.
A statistical batch is not just a memory setting
Processing 64 sources does not require keeping every encoder activation in memory simultaneously. A microbatch is a smaller group processed at one time. For objectives that couple examples, however, splitting the computation must preserve the intended statistics.
In this JEPA configuration, the projector’s BatchNorm sees 640 representations: ten views times 64 sources. VISReg is evaluated separately for each view across the 64 sources. These are two different statistical populations.
Calculating independent losses on small microbatches would generally change that objective. The implementation uses representation and gradient caching with encoder recomputation to retain the full-batch calculation while limiting activation memory. The important engineering question is whether the resulting gradients match the intended calculation, not merely whether the run fits on a GPU.
Connect the recipes to the measurements
The recorded feature spectra show that these choices lead to different representations: both matched-patch arms concentrate variation, full-field MAE spreads it relative to the initialization, and JEPA has the broadest spectrum in the measured panel. I discuss their checkpoint positions, interpretation, and the profiling behind the two-GPU implementation in Reading the geometry of AS-OCT pretraining. A common patient-separated downstream protocol is the next step for comparing task performance.
The raw losses also cannot serve as a leaderboard. MAE pixel error, diffusion reconstruction error and JEPA agreement with regularization have different targets and scales. Their numeric values are not a common unit of representation quality.
Try this: the number of anatomical candidates increases from 127 to 129. What happens to the matched budget?
At 127 candidates, M = 127, V = 32 and T = 95. At 129, the cap gives M = 128, V = 32 and T = 96. Both matched arms receive that same budget; neither gains an extra visible patch.
That small calculation captures the purpose of documenting the recipes precisely. A method name suggests a family of ideas. The experiment is the actual combination of data, views, targets, weighting, initialization and budget. Writing those choices down makes the work understandable, testable and worth comparing.
The first recorded representation diagnostics are discussed in Reading the geometry of AS-OCT pretraining, with dated curves, the shared initialization, and the profiling evidence.