WritingRepresentation learning

Learning without labels: what MAE and JEPA actually predict

Follow one image from pixels to representations, calculate a training loss, and see why agreement alone can teach a model nothing useful.

In this article

Cover part of a drawing with a sheet of paper. If the visible lines form an arc, you can make a reasonable guess about what continues behind the paper. Nobody needs to label the drawing before you can try this exercise: the answer is already underneath the cover.

That is a useful starting point for self-supervised learning. We construct a task from the data itself, train a model to solve it, and then ask whether the model has learned something useful beyond that task.

I implement these objectives for anterior-segment optical coherence tomography, or AS-OCT: cross-sectional images of the front of the eye. My experiments compare masked reconstruction with a global–local joint-embedding recipe. The different targets help explain the representation changes I measured.

We can understand the first piece with three synthetic drawings: a shallow arc, a steep arc and an asymmetric arc. Give each drawing two versions with slightly different brightness. There are no diseases or clinical labels in this example. We simply want to understand where the learning signal comes from.

First, follow the numbers

An image is an array of numbers. An encoder is a function that turns those numbers into another numerical representation. Its parameters determine how it performs that transformation, and training changes the parameters to reduce a chosen error.

An embedding is one such learned representation. Think of a vector such as [0.4, -1.2, 0.7], although real models often produce many more coordinates, sometimes one vector per image patch. Those coordinates do not arrive with names such as “curvature” or “disease.” Any useful structure has to emerge through learning and be investigated afterward.

Both masked autoencoders and joint-embedding methods learn representations. Their important difference is the target used to train them: what, exactly, is the prediction compared with?

MAE: use the missing pixels as the answer

In the masked autoencoder introduced by He and colleagues, the image is divided into patches. Some patches are hidden. The encoder processes the visible patches, and a decoder uses their representations and position information to predict the missing pixels. The decoder is the part that maps the learned representation back toward an image. The training error is calculated on the hidden content. MAE paper

On our synthetic arc, imagine that one hidden region contains just two values. The correct values are [0.2, 0.8]; the decoder predicts [0.3, 0.6]. The errors are 0.1 and -0.2. Squaring prevents the signs from cancelling, and averaging gives:

mean squared error = ((0.3 − 0.2)² + (0.6 − 0.8)²) / 2
                   = (0.01 + 0.04) / 2
                   = 0.025

The same operation extends to many hidden pixels. The optimizer adjusts the model so that future predictions reduce this error. When a recipe normalizes each target patch, the values being compared are those normalized values; the meaning of the loss depends on that preprocessing. The authors’ implementation includes this option. MAE implementation

Why might the encoder learn anything reusable? To predict the hidden part of an arc, it can benefit from knowing the visible contour, its position and the patterns that recur across examples. The representation is trained through a reconstruction task even if reconstruction is not the eventual application.

But the exercise has its own priorities. Reducing error on a large, predictable background may matter to the training objective. A small feature relevant to a future task may contribute very little. Whether the representation preserves that feature is an empirical question. A pleasing reconstruction is useful evidence about reconstruction.

Change the target: compare representations

Now take two views of our shallow arc. One might be a full drawing with a brightness change; the other might be a crop. Instead of asking a decoder to reproduce pixels, we can ask numerical representations of related content to agree.

This introduces a choice that pixel reconstruction largely avoids: the target representation is itself produced by a learned function. What produces it, and what prevents the system from taking an unhelpful shortcut?

“JEPA” names a family of approaches, so the answer depends on the method. I-JEPA, introduced by Assran and colleagues, predicts representations of target regions from a context region. Its target encoder is updated using an exponential moving average of the online encoder’s parameters. I-JEPA paper

LeJEPA, by Balestriero and LeCun, combines a predictive objective with explicit regularization called SIGReg. It does not use that teacher–student arrangement. A diagram containing a teacher is therefore a diagram of a particular method, not a definition of every JEPA. LeJEPA paper

For the rest of this explanation, I will use the shared-encoder JEPA+VISReg recipe recorded for my AS-OCT experiments in September 2026. This is a specific adaptation, distinct from both an unmodified I-JEPA and an unmodified LeJEPA recipe.

Our recipe: several views, one encoder

Each source image produces two global views and eight local views. Every view passes through the same encoder, followed by a projector: a small network that maps encoder features into the space where the training objective is applied.

For each source, the two global projections are averaged. All ten view projections are encouraged to approach that average. With a two-coordinate toy example, suppose the global projections are [1, 3] and [3, 1]. Their average is [2, 2]. A local projection [2, 3] has a mean squared difference of 0.5 from that target.

The real objective averages across coordinates, views and sources. Both global projections participate in the gradients. There is no separate teacher encoder and no stop-gradient applied to this agreement target.

This sounds promising: views of the same source should capture related information. But try solving the objective in the laziest possible way.

Give every view of every drawing the vector [2, 2].

The global average is [2, 2]. Every local vector matches it. The agreement error is zero. Yet the shallow arc, steep arc and asymmetric arc have become indistinguishable.

That is representation collapse. Zero is not special here. Any constant vector can create the same problem. A downstream classifier receiving only that constant cannot recover which drawing produced it, however impressive the agreement curve looks.

Perfect agreement can hide collapseSynthetic worked example
feature 1feature 2 1234

Paired-view error: 0. Relative spread: 1.00.

Each number is a different source, with two identical views. At zero separation, all sources share the same vector. These are constructed points, not a training run or a simulation of VISReg.

Move the control to spread the source vectors apart. The paired views still agree exactly, so their error stays at zero. This synthetic construction illustrates the ambiguity of agreement alone; it does not train a network or simulate the behavior of VISReg.

Agreement needs a second condition

The recipe therefore asks for two things at once: related views should agree, and the representations across different sources should retain a useful amount of variation. Making every image arbitrarily far from every other image would not express that requirement either.

VISReg, introduced by Wu, Balestriero and Levine, controls the scale of representations separately from their distributional shape. It uses projections onto one-dimensional directions to compare distributions. This differs from simply penalizing correlations between coordinates. VISReg paper

The local implementation makes that idea easier to inspect by exposing three components:

  • Center: penalize a nonzero average representation across sources.
  • Scale: penalize coordinate standard deviations that depart from the chosen unit scale.
  • Shape: center and standardize the representations, project them along sampled directions, then compare the sorted projections with Gaussian reference quantiles.

For an intuition about the last step, imagine viewing a cloud of points from several directions. Along each direction, the cloud becomes a list of positions on a line. These simpler views reveal aspects of its distribution that a single spread measurement misses. A finite set of projections is still a training construction; it is not a guarantee that every property of the full distribution is correct.

In the dated AS-OCT configuration, the total loss is 0.3 × agreement + 0.7 × regularization. Those coefficients describe this experiment. They are not the definition of JEPA or a recommendation for every dataset.

There is another important boundary. Regularization acts on the projected representations used for training. To understand what the encoder itself learned, I also need to inspect and evaluate encoder features. A tidy training space is not the whole result.

Diversity is necessary for this objective, but usefulness is another test

Return to our three drawings. Suppose the representations spread out beautifully, but mainly encode brightness. The model can distinguish examples, yet it may fail when the same arc appears under different illumination.

Conversely, suppose I ask the model to ignore any change that removes a small notch from an arc. If the future task is to detect that notch, I have encouraged the model to discard useful information. An invariance is a learned tolerance to a transformation. Choosing invariances means deciding which differences the representation should ignore.

This is why image augmentation deserves scientific attention. Cropping, reducing resolution and changing intensity determine which information remains available and which differences the objective rewards the model for overlooking. Their suitability depends on the task.

There are three distinct questions to check:

  1. Does optimization behave as intended: finite values, sensible gradients and reproducible state transitions?
  2. Does the representation retain variation, rather than converging to a constant or using very few directions?
  3. Does it help on an independently evaluated task that matters?

The first two help diagnose training. They cannot answer the third by themselves. Likewise, a lower MAE loss than JEPA loss has no direct ranking interpretation: the losses compare different objects, with different scales and normalizations.

A quick test of the idea

Question: every image produces [7, -3], and every augmented view matches its source perfectly. Has training succeeded?

Answer: the agreement objective has reached its best possible value, but the representation cannot distinguish the images. The vector being nonzero changes nothing. We would need to examine the diversity constraint and then evaluate useful information on a separate task.

Now transfer the question to audio. Two noisy versions of one recording could provide related views. But removing the frequencies that distinguish two speech sounds might erase exactly what a recognition system needs. Constructing a self-supervised task always embeds assumptions about which information matters.

In AS-OCT, these assumptions become choices about full-image context, native-resolution crops, anatomical masks and computation. The next article follows those choices through five recorded pretraining recipes. The useful starting point for reading any of them is the same: identify the input, identify the target, and ask what would let the model solve the exercise without learning what you hoped.