FFoveated: Spending Bits Where the Eye Is Looking
Steering an H.264 encoder with live eye-tracking data.
My MSc thesis asked a simple question: if you know exactly where someone is looking, how many bits can you stop sending everywhere else?
Only a Tiny Part of the Frame Is Actually Seen
The human retina is wildly non-uniform. Cone density peaks in the fovea centralis, right on the visual axis, and falls off steeply from there — outside a narrow region of roughly 2.5° around the current fixation point, colour and detail simply are not resolved at full acuity. Every video codec on the planet nevertheless encodes the whole frame as if the viewer were scrutinising all of it at once.
Region-of-interest coding tries to exploit this, but conventionally has to guess which region matters, using saliency models or face detectors. Those guesses are speculative and frequently wrong, which caps how aggressively you can degrade the rest of the frame. If you guess wrong and the viewer looks straight at the part you threw away, they notice immediately.
The alternative is not to guess. Put an eye-tracker on the viewer, feed the fixation point back to the encoder, and adapt the frame while they watch it. That is foveation, and the idea is old — Girod discussed it back in the late eighties — but he considered it impractical because of the round-trip delay between an eye movement and the updated image. Thirty years of improvements to trackers, encoders and networks have quietly made it feasible.
FFoveated
I built FFoveated, a framework for prototyping foveated coding schemes, on top of FFmpeg’s libav* libraries (hence the ajar name). The goal was never a novel codec, but a rig where the whole loop — gaze capture, encoding, playback, and the user study around it — could be swapped and re-parameterised quickly.

The awkward part is latency. A naive single-threaded decode-encode-decode cycle deadlocks itself on the very first frame, since all three codec instances want data nobody has produced yet. FFoveated instead splits reading, source decoding, foveated encoding, playback decoding and rendering across five threads joined by blocking FIFOs. The first two stages get generous 32-element buffers to absorb disk hiccups; the latency-critical stretch — from “encode this frame for where the eye is right now” to “show it” — gets buffers with a capacity of exactly one.
Fixation data reaches the encoder through FFmpeg’s side-data mechanism. I added a new AV_FRAME_DATA_FOVEATION_DESCRIPTOR type to AVFrameSideDataType carrying four floats: normalised x and y of the fixation point, the standard deviation σ, and the maximum quality offset δ. Nicely, this is the only codec-independent patch needed — and if you link against unpatched libav*, the wrappers just free the unknown side data and you get ordinary unfoveated encoding for free.
The Offset Function
The actual foveation happens in the x264 wrapper. H.264’s quantization parameter runs from 0 to 51, with higher values meaning coarser quantization and fewer bits. Rather than a hard region-of-interest rectangle, each macroblock gets a smooth offset added to its qp:

An inverted Gaussian is the natural shape given the circular fovea, and it has the nice property of inflicting no penalty at the fixation point itself. I fixed σ at 2.5° of visual angle, straight from the retinal characteristics — an educated approximation rather than a tuned optimum, since the mapping from qp to perceived quality depends heavily on content. That leaves δ, the peripheral penalty, as the single knob to push: how far can you crank it before someone notices?
For the codec configuration itself, real-time constraints dictate most choices: the ultrafast preset with zerolatency tuning (which, among other things, disables B-frames, since predicting from future frames is impossible on a causal source), a GOP limit of three frames to stop quantization error accumulating and to survive packet loss, and aq-mode 1 so there is a variance-based adaptive quantization pass to add the offset into. The output remains standard-conforming H.264 — any normal player will decode it.
Finding the Point Where People Notice
Rather than collecting mean opinion scores, the study hunted for each viewer’s just noticeable distortion threshold directly. Ten participants, ten source videos from the VQEG JEG Hybrid dataset (10 seconds each, 1080p at 25 fps), presented on a colour-calibrated UHD screen with an SMI RED250mobile tracker mounted below it.

Each source was shown ten times in a row. Within a repetition, δ ratchets upward every few frames, so quality in the periphery decays continuously as you watch. The moment a participant sees an artefact they press a button; the repetition stops, a black screen flashes for a second, δ drops back by a set decrement, and the next repetition begins with a gentler ramp. Early repetitions climb fast and fall far, later ones creep — a repeated linear search converging on that person’s threshold for that particular clip. It also sidesteps the usual problem with subjective testing: nobody has to hold an abstract five-point quality scale in their head, they just have to say “there, I saw it.”
That produced 1000 video presentations and 734 interaction events.
63% Fewer Bits
The headline result: at the 10% JND — the strict definition, where only one in ten viewers reports visible distortion — foveated encoding used 62.76% fewer bits than the identical encoder with δ = 0. Averaged across the ten sources, 6980 kbit/s fell to 2652 kbit/s. At the more commonly used 25% JND the figure rises to 68.88%, though that number is flattered by the aggressive δ ramp early in each sequence.

The frame above is the whole thesis in one image. All the action is confined to the left side of the court, the viewer is locked onto it, and the stadium ranks — which occupy most of the pixels — have been quantized into mush that nobody looked at closely enough to notice. Sports content like this is close to the ideal case; the savings are content-dependent, ranging from 54% on the busiest clip to 71% on the most forgiving one.
For context, the comparable numbers in the literature at the time were around 63% (Arndt and Antons) and around 42% (Illahi et al., for cloud-rendered game streaming). Direct comparison is genuinely difficult, since anchoring on the JND was novel and the field had no common benchmark — but the picture is consistent.
What the Eye-Tracker Told Me About the Participants
The most unexpectedly interesting data was not the bitrates but the gaze paths. Compare two participants watching the same clip — a cheetah pacing back and forth:


The first is exactly what you would hope for: a tight band of fixations following the region of interest. The second is a participant who worked out what the experiment was measuring and went hunting for artefacts in the periphery — large saccades terminating repeatedly in objectively uninteresting background regions. That is not natural viewing behaviour, and it systematically drags their reported threshold down.
Which suggests something reusable beyond this study: eye-tracking data could serve as a post-hoc filter on participants in quality assessment databases, separating people who watched from people who audited. As far as I know that had not been done in image or video quality assessment. I wrote that idea up separately, along with the observation that foveated coding gives you a strong prior on where a visible distortion must be — which in principle lets you infer noticeability from gaze irregularities alone, without asking anyone to press a button at all. The dataset here was too small to train anything on it, but the hook is there.
Where It Goes
Bigger screens and higher resolutions should push the savings further, since a larger share of every frame lands in the periphery — with encoding time as the likely bottleneck. Peripheral vision is disproportionately sensitive to abrupt contrast changes, so blurring the coarsely quantized blocks would probably buy additional headroom. And the lab dependency is the real barrier to scale: approximate webcam-based eye-tracking would be plenty precise for something this spatially forgiving, and would put the whole approach within reach of crowdsourced studies and, more to the point, of ordinary video calls.
The full thesis is available as a PDF. Parts of it were published as Foveated Video Coding for Real Time Streaming Applications at QoMEX 2020, and the participant-filtering ideas as Gaze Data for Quality Assessment of Foveated Video at the ET-MM workshop at ACM ETRA 2020.