<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom"><title>Oliver Gmelin</title><link href="https://gmelin.io/" rel="alternate"/><link href="https://gmelin.io/feeds/all.atom.xml" rel="self"/><id>https://gmelin.io/</id><updated>2026-09-08T19:09:51.876540+02:00</updated><entry><title>The July 30 pressure wave / 100km Alpstein FAI</title><link href="https://gmelin.io/blog/alpstein-fai/" rel="alternate"/><published>2026-07-30T00:00:00+02:00</published><updated>2026-09-08T19:09:51.876540+02:00</updated><author><name>Oliver Gmelin</name></author><id>tag:gmelin.io,2026-07-30:/blog/alpstein-fai/</id><summary type="html">&lt;p&gt;Risk mitigation: A homage to modern (in-)flight information systems.&lt;/p&gt;</summary><content type="html">&lt;p&gt;&lt;strong&gt;Risk mitigation: A homage to modern (in-)flight information systems.&lt;/strong&gt;&lt;br&gt;
Technology is often considered &lt;em&gt;unromantic&lt;/em&gt; in mountain sports, but some hassles are well worth it.&lt;/p&gt;
&lt;h2&gt;Tech Aversion in Mountain Sports&lt;/h2&gt;
&lt;p&gt;I am under the impression that there&amp;rsquo;s still a rift in the mountain sports and paragliding communities regarding the usage of technology. Devices and resources are sometimes waved aside as unnecessary, obtrusive or unromantic. After all, LCD screens distract us from true adventure and genuine experiences in nature, right? Mountaineers and pilots thereby miss out on safety enhancing, possibly life-saving technologies.&lt;/p&gt;
&lt;p&gt;On July 30 2026, large parts of Europe were affected by a dry air pressure wave. Violent wind gusts originated from an array of thunderstorms cells over France and swept eastwards without any warning signs: No clouds, no steady increase in wind speeds, just life-threatening air conditions from one second to the next. This rare occurrence was predicted by the ICON models, albeit the exact timing remained a bit uncertain.  It caused a stir on social media, most of the pertinent weather apps and services posted warnings about the following day.&lt;/p&gt;
&lt;figure class="w-full"&gt;
&lt;p&gt;&lt;img alt="passing the Saentis" src="https://gmelin.io/blog/alpstein-fai/pressurewave.gif"&gt;&lt;/p&gt;
&lt;figcaption&gt;Excellent depiction of the incoming threat in &lt;a href="https://burnair.cloud"&gt;Burnair&lt;/a&gt;.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;strong&gt;Takeaway 1:&lt;/strong&gt; Do an exhaustive meteo briefing. Every single time. No exemptions.&lt;/p&gt;
&lt;p&gt;Days like this prove that you will never outsmart modern weather models even with years of experience on your local site. There&amp;rsquo;s a nonzero chance that things turn bad in ways you won&amp;rsquo;t see coming. Depending on your planned activity, proper &amp;ldquo;state of the art&amp;rdquo; planning preparation likely requires a (paid) subscription to a specialized weather service. That&amp;rsquo;s a small price in comparison to the potential cost of an accident.&lt;/p&gt;
&lt;p&gt;Thorough weather briefings are already a point of contention. People sometimes justify subpar preparation because their plans are allegedly a special case. &amp;ldquo;It&amp;rsquo;s just a quick hike and fly!&amp;rdquo; Sure, conditions the last 30 minutes might have been calm, but that could have led to a severe accident on July 30. Preparatory weather checks are mandatory whether you&amp;rsquo;re hiking, mountaineering, crosscountry- or speedflying. Only if you know exactly what&amp;rsquo;s coming, you can manage risks with adequate safety margins. This, on the other hand, opens up potentials on less than ideal days.&lt;/p&gt;
&lt;h2&gt;In-Flight Hassle vs Reward&lt;/h2&gt;
&lt;p&gt;While solid preparation is out of the question, I&amp;rsquo;d actually distinguish the recommended PG tech based on the planned activity. Short glides or even speedflying won&amp;rsquo;t give you much chance to check for weather updates. The measurement intervals are just too long, you&amp;rsquo;re busy flying and local indicators (trees, windsocks etc.) are just better scoped.&lt;/p&gt;
&lt;p&gt;As soon as there&amp;rsquo;s a chance to thermal upwards or even dip into XCish terrain, I&amp;rsquo;d argue that in-flight weather information systems become sensible. Looking at the animation above, we can see that the measured wind speeds (isolated arrows) were actually increasing ahead of the predicted pressure wave. While modern weather predictions are pretty good, &lt;strong&gt;live weather updates still provide significant safety benefits&lt;/strong&gt;. As depicted below, there were really no warning signs about the incoming danger.&lt;/p&gt;
&lt;figure class="figure-grid" style="--figure-columns: 1fr 1.745fr 1fr"&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="passing the Saentis" src="https://gmelin.io/blog/alpstein-fai/saentis.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;About to pass by the Saentis.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="Ridge" src="https://gmelin.io/blog/alpstein-fai/ridge.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;Blue skies all around.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="Walensee" src="https://gmelin.io/blog/alpstein-fai/walensee.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;En route along Walensee.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;strong&gt;Takeaway 2:&lt;/strong&gt; Anything that could turn into a longer thermal or XC flight warrants in-flight weather data.&lt;/p&gt;
&lt;h2&gt;Collision Avoidance Systems&lt;/h2&gt;
&lt;p&gt;Collision avoidance systems based on FLARM and the newer ADS-L systems are not as widely adopted as one would expect in 2026. Granted, between paragliders they&amp;rsquo;re often useless: when thermalling together you just have to stay vigilant and collisions in low-visibility conditions (e.g. being sucked in a cloud) should not happen in the first place. But paragliders share the airspace with faster aircraft, whether it be sailplanes, motorized sporting planes, helicopters or even low-flying military jets.&lt;/p&gt;
&lt;p&gt;A spectacular &lt;a href="https://aviation-safety.net/wikibase/570859"&gt;collision between a paraglider and Cessna&lt;/a&gt; in May shows that this is not a purely theoretical concern. If you&amp;rsquo;ve ever wondered whether &lt;em&gt;that sail plane over there&lt;/em&gt; had actually seen you or when jets in the clouds above sounded too close for comfort you probably pondered on getting a vario with FLARM/ADS-L. Hopefully you acted on that thought. Caveat emptor: Currently there&amp;rsquo;s no guarantee other aircraft will actually use and see you on the necessary receivers.
There&amp;rsquo;s rumors regarding mandatory electronical identification measures for paragliders in Switzerland. This has the potential to replace the unpopular stick-on labels. For U spaces that will be shared with autonomously flying drones, there will be no way around ADS-L anyways.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Takeaway 3:&lt;/strong&gt; At least for thermalling/XC, get an FLARM/ADS-L transceiver.&lt;/p&gt;
&lt;h2&gt;Online Tracking and Distress Calls&lt;/h2&gt;
&lt;p&gt;The last puzzle piece of in-flight electronics/services that comes to mind is online tracking. If you&amp;rsquo;re using FLARM/ADS-L this is potentially covered through &lt;a href="http://glidernet.org/"&gt;OGN&lt;/a&gt; receivers already, depending on your region. Alternative are GSM-based, e.g. the &lt;a href="https://www.rega.ch/en/our-missions/this-is-how-we-help-you/rega-app"&gt;rega&lt;/a&gt; app or &lt;a href="https://burnair.cloud"&gt;Burnair&lt;/a&gt;. These all rely (more or less) on a line of sight to a corresponding cellular transmitter. That&amp;rsquo;s already an enormous benefit: the search radius for a rescue operation can be narrowed down significantly if your last airborne position is known. There&amp;rsquo;s simply no excuse to take off without app-based live tracking enabled.&lt;/p&gt;
&lt;p&gt;The last failure mode is rooted in the  unavailability of terrestrial radio networks. In the air, cell coverage in the Alps is pretty good. In the event of a reserve deployment and a rough landing somewhere on a glacier you have a good chance to be found at some point if you had a GSM tracker running and someone notices that you&amp;rsquo;re out way longer than anticipated. The key advantage is time: being able to call for help immediately even without network coverage might be a life-saver and cold nights can turn threatening. &lt;/p&gt;
&lt;p&gt;There&amp;rsquo;s multiple vendors for satellite communication that bridge the last remaining gap. The Iridium and Globalstar networks have global coverage, others are more region specific. Pick carefully, depending on your location. For adventures in remote areas these are a must-have.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Takeaway 4:&lt;/strong&gt; Get a satellite communicator as soon as you start solo XC flying.&lt;/p&gt;
&lt;h2&gt;Common Counterarguments&lt;/h2&gt;
&lt;p&gt;I think a major point essentially boils down to &lt;em&gt;not wanting to deal with yet another complicated piece of equipment&lt;/em&gt;, i.e. laziness. We should overcome this as a community. Another counterargument is the price point: Granted, a complete setup isn&amp;rsquo;t cheap. But considering the possible consequences of a paragliding accident I think it&amp;rsquo;s well invested and not paying off on a personal level, but also for the entire community.&lt;/p&gt;
&lt;p&gt;The safer paragliding is, the less severe the consequences weather forecast misinterpretations and accidents become, the fewer ramifications and resultant regulations will we see as a community. Maybe we should adapt the &amp;ldquo;all the gear, all the time&amp;rdquo; credo from motorcycling.&lt;/p&gt;
&lt;h2&gt;The 100km FAI Triangle on July 30&lt;/h2&gt;
&lt;p&gt;Despite the pressure wave prediction I went flying on July 30. Absolutely fantastic day. Would have been nice to fly to Lucerne or Interlaken, but I obviously turned around and opted for the 100km triangle instead. The pressure wave was clearly visible on live weather in-flight.&lt;/p&gt;
&lt;figure class="w-wide" style="--figure-width: 30rem"&gt;
&lt;p&gt;&lt;img alt="Tracklog preview." src="https://gmelin.io/blog/alpstein-fai/map-cropped.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;Tracklog preview.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Full logs are available on &lt;a href="https://en.dhv-xc.de/flight/2280506"&gt;DHV XC&lt;/a&gt;, &lt;a href="https://www.xcontest.org/world/en/flights/detail:ogmln/30.7.2026/09:56"&gt;xcontest&lt;/a&gt; and &lt;a href="https://igcviewer.burnair.cloud/?file=https://burnair-activities-prod.s3.eu-central-1.amazonaws.com/11-52CC/2026-07-30_11.56_EBENALP_NW.IGC"&gt;Burnair 3D&lt;/a&gt;&lt;/p&gt;</content><category term="paragliding"/><category term="flights"/></entry><entry><title>Computer Vision for Aquatic Biodiversity Exploratories</title><link href="https://gmelin.io/blog/project-above-updates/" rel="alternate"/><published>2026-06-20T00:00:00+02:00</published><updated>2026-09-08T19:09:33.889026+02:00</updated><author><name>Oliver Gmelin</name></author><id>tag:gmelin.io,2026-06-20:/blog/project-above-updates/</id><summary type="html">&lt;p&gt;Implementing imaging systems for automated long-term experiments on aquatic evolution&lt;/p&gt;</summary><content type="html">&lt;p&gt;&lt;strong&gt;Implementing imaging systems for automated long-term experiments on aquatic evolution.&lt;/strong&gt;&lt;br&gt;
A sporadically updated project log.&lt;/p&gt;
&lt;h2&gt;Intro&lt;/h2&gt;
&lt;p&gt;Biodiversity is declining and freshwater lakes are not exempt from this global trend.
Modeling and understanding these huge systems is nontrivial.
Classical lab setups, e.g. chemostats, do not provide sufficient volume and operational lifetime to
observe the intricate interactions and derive robust underlying interaction patterns between species and environmental conditions.&lt;/p&gt;
&lt;figure class="figure-grid grid-narrow"&gt;
&lt;figcaption&gt;&lt;strong&gt;Small-volume chemostats: insufficient models for entire lakes&lt;/strong&gt;&lt;/figcaption&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="chemostats fresh" src="https://gmelin.io/blog/project-above-updates/chemostats-stabilized-2s.gif"&gt;&lt;/p&gt;
&lt;figcaption&gt;Freshly initialized chemostats.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="chemostats green" src="https://gmelin.io/blog/project-above-updates/chemostat_algae.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;Algal growth.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/figure&gt;
&lt;h2&gt;Mesocosm Deployments&lt;/h2&gt;
&lt;p&gt;With project &lt;strong&gt;ABOVE&lt;/strong&gt;, we aim for (mostly) automated long-term experiments on the evolution of large-scale aquatic model systems. Multiple sets of &lt;em&gt;mesocosms&lt;/em&gt;, which are 600 liter tanks, are set up in groups of four. 
These share a multiparameter probe for chemical properties as well as flow cytometers for plankton imaging. &lt;/p&gt;
&lt;figure class="figure-grid" style="--figure-columns: 1.27fr 1fr"&gt;
&lt;figcaption&gt;&lt;strong&gt;ABOVE mesocosm deployments&lt;/strong&gt;&lt;/figcaption&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="mesocosm tankgroup" src="https://gmelin.io/blog/project-above-updates/meso_tankgroup.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;revised tank group installation&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="control infrastructure deployment" src="https://gmelin.io/blog/project-above-updates/control_deployment.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;provisional control infrastructure&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/figure&gt;
&lt;p&gt;The depicted white tanks are already an improved, winter-hardened version. The tanks are not hermetically sealed, but allow for a reasonable environmental exchange (e.g. regarding oxygen) similar to that of a real lake. Internal flow pumps prevent the water from turning stale, a refill mechanism compensates for sample extraction and the temperatures are controlled within a viable range. Lights allow controlling the algal energy intake. The tanks themselves, the fluid handling and the control systems were implemented and manufactured by &lt;a href="https://www.4h-jena.de/en/maritime-technologies/mesocosms/"&gt;4H Jena Engineering&lt;/a&gt;.&lt;/p&gt;
&lt;figure class="figure-grid" style="--figure-columns: 1fr 1.34fr"&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="old mesocosms filled" src="https://gmelin.io/blog/project-above-updates/meso_old_filled.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;Test runs in the V1 tanks.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="new tank group wrapped" src="https://gmelin.io/blog/project-above-updates/tankgroup_wrapped.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;Tank delivery on the new platform.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/figure&gt;
&lt;h2&gt;Analysis Automation&lt;/h2&gt;
&lt;p&gt;With intended operational times in the ballpark of decades and many replicates, experiments of this extent would constitute a significant expenditure of human labor.
This is where the automation aspect comes into play: flow cytometry instead of manual sampling and traditional microscopy, continuous 24/7 monitoring of relevant parameters
and fusion of comprehensive time-series data is expected to generate insights into the evolution of aquatic systems.&lt;/p&gt;
&lt;figure class="figure-grid" style="--figure-columns: 2.5fr 1fr"&gt;
&lt;figcaption&gt;&lt;strong&gt;Computer vision to the rescue&lt;/strong&gt;&lt;/figcaption&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="lab culture segmentation" src="https://gmelin.io/blog/project-above-updates/segmentation.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;Flow cytometry image and instance segmentation.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure&gt;
&lt;video autoplay loop muted playsinline src="https://gmelin.io/blog/project-above-updates/crops-10fps.mp4"&gt;&lt;/video&gt;
&lt;figcaption&gt;Segmented objects at 10fps.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/figure&gt;
&lt;p&gt;To this extent, we have provisioned several hundred terabytes of storage to keep raw data and derived descriptors for the foreseeable future.  Results will follow. Probably as a paper, not so much a blog post. &lt;/p&gt;</content><category term="science"/><category term="computervision"/><category term="devops"/></entry><entry><title>LLMs on a GP102/Pascal 3-Way SLI Rig</title><link href="https://gmelin.io/blog/pascal-ollama/" rel="alternate"/><published>2025-12-20T00:00:00+01:00</published><updated>2026-09-01T17:05:12.889609+02:00</updated><author><name>Oliver Gmelin</name></author><id>tag:gmelin.io,2025-12-20:/blog/pascal-ollama/</id><summary type="html">&lt;p&gt;Hassle-free local inference through Vulkan&lt;/p&gt;</summary><content type="html">&lt;p&gt;&lt;strong&gt;Hassle-free local inference through Vulkan.&lt;/strong&gt;&lt;br&gt;
Nvidia keeps nudging Pascal further out the door, but there&amp;rsquo;s an easier way to keep these cards doing useful work than fighting the CUDA toolchain on a rolling distro.&lt;/p&gt;
&lt;h2&gt;Three Cards, One Dead Standard&lt;/h2&gt;
&lt;p&gt;The rig is a leftover from when multi-GPU gaming was still a thing people did: three GP102-based GeForce cards on an X299 board, bridged for 3-way SLI, fed by a 1600W supply that has spent most of the last few years being wildly overqualified for the job.&lt;/p&gt;
&lt;figure class="w-wide"&gt;
&lt;p&gt;&lt;img alt="Three GeForce GTX cards in a SLI setup" src="https://gmelin.io/blog/pascal-ollama/gpus.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;Three-way GP102 setup.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;SLI itself is thoroughly dead — no driver profile is coming to make three cards render one game any more. But that has nothing to do with whether the chips can still compute. Each GP102 is a perfectly capable device with a useful amount of VRAM attached, and three of them idling next to an oversized PSU is a waste when running models locally has become the normal way to try anything out. So: &lt;a href="https://ollama.com/"&gt;ollama&lt;/a&gt;, and how hard can it be.&lt;/p&gt;
&lt;h2&gt;Nvidia Shows Pascal the Door&lt;/h2&gt;
&lt;p&gt;Harder than it should be, because Pascal is now actively being removed from the stack — twice over, at two different layers.&lt;/p&gt;
&lt;p&gt;The first wall is the driver. Nvidia&amp;rsquo;s 590 branch &lt;a href="https://archlinux.org/news/nvidia-590-driver-drops-pascal-support-main-packages-switch-to-open-kernel-modules/"&gt;drops support for Pascal and everything older&lt;/a&gt;, as part of the same move that makes the open-source kernel modules the default. Pascal predates the GSP firmware those modules rely on, so it cannot follow. On Arch this is not a gentle deprecation: the &lt;code&gt;nvidia&lt;/code&gt;, &lt;code&gt;nvidia-dkms&lt;/code&gt; and &lt;code&gt;nvidia-lts&lt;/code&gt; packages roll forward to 590, and if you let that happen on a Pascal box you get a broken graphical environment. The prescribed fix is to uninstall those and install &lt;code&gt;nvidia-580xx-dkms&lt;/code&gt; from the AUR instead.&lt;/p&gt;
&lt;p&gt;The good news is that 580 is an LTS branch, and Nvidia has committed to supporting pre-Turing hardware on it into mid-2028. So the card keeps working — you just pin it and stop upgrading that one thing.&lt;/p&gt;
&lt;h2&gt;CUDA 13 Can&amp;rsquo;t Build for It Either&lt;/h2&gt;
&lt;p&gt;The second wall is the toolkit, and it is the one that actually bites. CUDA 13.0 removed &lt;strong&gt;offline compilation for everything below compute capability 7.5&lt;/strong&gt; — which sweeps up Maxwell, Pascal and Volta together. GP102 is &lt;code&gt;sm_61&lt;/code&gt;. &lt;code&gt;nvcc&lt;/code&gt; from CUDA 13 will not emit device code for it at all; there is no flag to talk it round, because the target support is gone rather than merely deprecated.&lt;/p&gt;
&lt;p&gt;Worth being precise about what this does and does not break. It is a &lt;em&gt;build-time&lt;/em&gt; wall, not a runtime one: binaries compiled with CUDA 12.9 or earlier keep running fine on any hardware the installed driver supports. Nothing already working stops working. But the moment you want to compile something yourself — which is exactly what you are doing if you want an ollama build that targets your cards — the last toolkit that can do it is 12.9.&lt;/p&gt;
&lt;h2&gt;The Toolchain Rabbit Hole&lt;/h2&gt;
&lt;p&gt;Pinning an old toolkit sounds simple. On a rolling-release distro it starts a chain.&lt;/p&gt;
&lt;p&gt;CUDA 12.9 supports GCC up to &lt;strong&gt;14.x&lt;/strong&gt; and no further. Arch has long since moved past that, so the system compiler is rejected outright — not with a subtle miscompile, but with a hard version check baked into CUDA&amp;rsquo;s &lt;code&gt;host_config.h&lt;/code&gt;:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;unsupported GNU version! gcc versions later than 14 are not supported!
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;

&lt;p&gt;So you install &lt;code&gt;gcc14&lt;/code&gt; alongside your real compiler and point &lt;code&gt;nvcc&lt;/code&gt; at it with &lt;code&gt;-ccbin&lt;/code&gt;, or wave the check away with &lt;code&gt;-allow-unsupported-compiler&lt;/code&gt; and hope. Getting either of those threaded through a package build&amp;rsquo;s CMake configuration is its own small adventure.&lt;/p&gt;
&lt;p&gt;Clear that, and the next layer surfaces: an old CUDA toolkit&amp;rsquo;s headers against a current libstdc++ don&amp;rsquo;t agree on exception specifications any more, and the build dies on mismatched declarations. Patching &lt;code&gt;noexcept(true)&lt;/code&gt; onto the offending declarations in CUDA&amp;rsquo;s own headers gets you moving again.&lt;/p&gt;
&lt;p&gt;Add it up and the &amp;ldquo;working&amp;rdquo; configuration is a pinned driver, a pinned toolkit, a second compiler installed solely to appease the first one, and hand-edited vendor headers — a stack of four things that each break independently on the next update. That is not a setup, that is a pet.&lt;/p&gt;
&lt;h2&gt;Vulkan Sidesteps All of It&lt;/h2&gt;
&lt;p&gt;None of it is necessary, because CUDA is not the only way to get compute out of these cards.&lt;/p&gt;
&lt;p&gt;llama.cpp&amp;rsquo;s Vulkan backend — which ollama can be built against — talks to the GPU through the Vulkan compute API, and its shaders are compiled at runtime by the driver. There is no offline device-code generation step, so there is no &lt;code&gt;nvcc&lt;/code&gt;, no architecture flag, no host compiler ceiling and no vendor headers to patch. The 580xx driver you are already pinned to for display purposes ships a perfectly good Vulkan implementation. What got cut for Pascal was the CUDA build toolchain, not the silicon&amp;rsquo;s ability to run compute shaders.&lt;/p&gt;
&lt;p&gt;In practice this means swapping &lt;code&gt;vulkan&lt;/code&gt; in for the &lt;code&gt;cuda_v*&lt;/code&gt; entries in &lt;code&gt;OLLAMA_LLAMA_BACKENDS&lt;/code&gt; when building the package, and that is essentially the whole intervention.&lt;/p&gt;
&lt;p&gt;The trade-off is real but narrow. Vulkan and CUDA land in roughly the same place on token generation, which is the number you feel when a model is answering you. Prompt processing is where CUDA still has a clear lead, so long-context work and big prompt ingests are noticeably slower. For interactive use on hardware this old, that is a good trade.&lt;/p&gt;
&lt;h2&gt;Worth It?&lt;/h2&gt;
&lt;p&gt;For three cards that were otherwise going to sit in a cupboard, comfortably. The honest framing is that Pascal has entered its long tail: the driver is pinned until 2028, the CUDA path is closed to new builds, and nothing about that is going to improve. Vulkan is the option that doesn&amp;rsquo;t fight it — it asks nothing of the toolchain, so there is nothing in the setup to break the next time something upstream moves.&lt;/p&gt;</content><category term="science"/><category term="LLMs"/></entry><entry><title>FFoveated: Spending Bits Where the Eye Is Looking</title><link href="https://gmelin.io/blog/ffoveated/" rel="alternate"/><published>2020-06-11T00:00:00+02:00</published><updated>2026-09-01T16:25:01.276812+02:00</updated><author><name>Oliver Gmelin</name></author><id>tag:gmelin.io,2020-06-11:/blog/ffoveated/</id><summary type="html">&lt;p&gt;Steering an H.264 encoder with live eye-tracking data cut bitrates by 63% at the point where viewers first noticed the difference.&lt;/p&gt;</summary><content type="html">&lt;p&gt;&lt;strong&gt;Steering an H.264 encoder with live eye-tracking data.&lt;/strong&gt;&lt;br&gt;
My MSc thesis asked a simple question: if you know exactly where someone is looking, how many bits can you stop sending everywhere else?&lt;/p&gt;
&lt;h2&gt;Only a Tiny Part of the Frame Is Actually Seen&lt;/h2&gt;
&lt;p&gt;The human retina is wildly non-uniform. Cone density peaks in the fovea centralis, right on the visual axis, and falls off steeply from there — outside a narrow region of roughly 2.5° around the current fixation point, colour and detail simply are not resolved at full acuity. Every video codec on the planet nevertheless encodes the whole frame as if the viewer were scrutinising all of it at once.&lt;/p&gt;
&lt;p&gt;Region-of-interest coding tries to exploit this, but conventionally has to &lt;em&gt;guess&lt;/em&gt; which region matters, using saliency models or face detectors. Those guesses are speculative and frequently wrong, which caps how aggressively you can degrade the rest of the frame. If you guess wrong and the viewer looks straight at the part you threw away, they notice immediately.&lt;/p&gt;
&lt;p&gt;The alternative is not to guess. Put an eye-tracker on the viewer, feed the fixation point back to the encoder, and adapt the frame &lt;em&gt;while they watch it&lt;/em&gt;. That is foveation, and the idea is old — Girod discussed it back in the late eighties — but he considered it impractical because of the round-trip delay between an eye movement and the updated image. Thirty years of improvements to trackers, encoders and networks have quietly made it feasible.&lt;/p&gt;
&lt;h2&gt;FFoveated&lt;/h2&gt;
&lt;p&gt;I built &lt;strong&gt;FFoveated&lt;/strong&gt;, a framework for prototyping foveated coding schemes, on top of FFmpeg&amp;rsquo;s &lt;code&gt;libav*&lt;/code&gt; libraries (hence the ajar name). The goal was never a novel codec, but a rig where the whole loop — gaze capture, encoding, playback, and the user study around it — could be swapped and re-parameterised quickly.&lt;/p&gt;
&lt;figure class="w-full"&gt;
&lt;p&gt;&lt;img alt="Block diagram: source decoder feeding a foveated encoder, which streams to a client-side decoder and display, with an eye tracker feeding fixation data back to the encoder" src="https://gmelin.io/blog/ffoveated/architecture.png"&gt;&lt;/p&gt;
&lt;figcaption&gt;A source (file or live) is decoded, re-encoded with foveation, and played back. The eye tracker closes the loop straight back into the encoder.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The awkward part is latency. A naive single-threaded decode-encode-decode cycle deadlocks itself on the very first frame, since all three codec instances want data nobody has produced yet. FFoveated instead splits reading, source decoding, foveated encoding, playback decoding and rendering across five threads joined by blocking FIFOs. The first two stages get generous 32-element buffers to absorb disk hiccups; the latency-critical stretch — from &amp;ldquo;encode this frame for where the eye is &lt;em&gt;right now&lt;/em&gt;&amp;rdquo; to &amp;ldquo;show it&amp;rdquo; — gets buffers with a capacity of exactly one.&lt;/p&gt;
&lt;p&gt;Fixation data reaches the encoder through FFmpeg&amp;rsquo;s side-data mechanism. I added a new &lt;code&gt;AV_FRAME_DATA_FOVEATION_DESCRIPTOR&lt;/code&gt; type to &lt;code&gt;AVFrameSideDataType&lt;/code&gt; carrying four floats: normalised x and y of the fixation point, the standard deviation σ, and the maximum quality offset δ. Nicely, this is the &lt;em&gt;only&lt;/em&gt; codec-independent patch needed — and if you link against unpatched &lt;code&gt;libav*&lt;/code&gt;, the wrappers just free the unknown side data and you get ordinary unfoveated encoding for free.&lt;/p&gt;
&lt;h2&gt;The Offset Function&lt;/h2&gt;
&lt;p&gt;The actual foveation happens in the x264 wrapper. H.264&amp;rsquo;s quantization parameter runs from 0 to 51, with higher values meaning coarser quantization and fewer bits. Rather than a hard region-of-interest rectangle, each macroblock gets a smooth offset added to its qp:&lt;/p&gt;
&lt;figure class="w-wide"&gt;
&lt;p&gt;&lt;img alt="3D plot of the quality offset function: a Gaussian well dipping to zero at the fixation point and flattening out at delta everywhere else" src="https://gmelin.io/blog/ffoveated/offset-function.png"&gt;&lt;/p&gt;
&lt;figcaption&gt;The offset is an inverted 2D Gaussian: zero penalty at the fixation point, saturating at δ in the periphery.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;An inverted Gaussian is the natural shape given the circular fovea, and it has the nice property of inflicting &lt;em&gt;no&lt;/em&gt; penalty at the fixation point itself. I fixed σ at 2.5° of visual angle, straight from the retinal characteristics — an educated approximation rather than a tuned optimum, since the mapping from qp to perceived quality depends heavily on content. That leaves δ, the peripheral penalty, as the single knob to push: how far can you crank it before someone notices?&lt;/p&gt;
&lt;p&gt;For the codec configuration itself, real-time constraints dictate most choices: the &lt;code&gt;ultrafast&lt;/code&gt; preset with &lt;code&gt;zerolatency&lt;/code&gt; tuning (which, among other things, disables B-frames, since predicting from future frames is impossible on a causal source), a GOP limit of three frames to stop quantization error accumulating and to survive packet loss, and &lt;code&gt;aq-mode 1&lt;/code&gt; so there is a variance-based adaptive quantization pass to add the offset into. The output remains standard-conforming H.264 — any normal player will decode it.&lt;/p&gt;
&lt;h2&gt;Finding the Point Where People Notice&lt;/h2&gt;
&lt;p&gt;Rather than collecting mean opinion scores, the study hunted for each viewer&amp;rsquo;s &lt;strong&gt;just noticeable distortion&lt;/strong&gt; threshold directly. Ten participants, ten source videos from the VQEG JEG Hybrid dataset (10 seconds each, 1080p at 25 fps), presented on a colour-calibrated UHD screen with an SMI RED250mobile tracker mounted below it.&lt;/p&gt;
&lt;figure class="w-full"&gt;
&lt;p&gt;&lt;img alt="Grid of ten video thumbnails: a cartoon, two basketball scenes, a cheetah, a lion, colourful toys, a lab, an aerial park view, zebras, and an escalator hall" src="https://gmelin.io/blog/ffoveated/vqeg-sources.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;The ten VQEG JEG sources — deliberately diverse: flat cartoon areas, fast action in small regions, high dynamic range, scene cuts and camera motion.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Each source was shown ten times in a row. Within a repetition, δ ratchets upward every few frames, so quality in the periphery decays continuously as you watch. The moment a participant sees an artefact they press a button; the repetition stops, a black screen flashes for a second, δ drops back by a set decrement, and the next repetition begins with a gentler ramp. Early repetitions climb fast and fall far, later ones creep — a repeated linear search converging on that person&amp;rsquo;s threshold for that particular clip. It also sidesteps the usual problem with subjective testing: nobody has to hold an abstract five-point quality scale in their head, they just have to say &amp;ldquo;there, I saw it.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;That produced 1000 video presentations and 734 interaction events.&lt;/p&gt;
&lt;h2&gt;63% Fewer Bits&lt;/h2&gt;
&lt;p&gt;The headline result: at the 10% JND — the strict definition, where only one in ten viewers reports visible distortion — foveated encoding used &lt;strong&gt;62.76% fewer bits&lt;/strong&gt; than the identical encoder with δ = 0. Averaged across the ten sources, 6980 kbit/s fell to 2652 kbit/s. At the more commonly used 25% JND the figure rises to 68.88%, though that number is flattered by the aggressive δ ramp early in each sequence.&lt;/p&gt;
&lt;figure class="w-wide"&gt;
&lt;p&gt;&lt;img alt="A basketball frame with sharp detail around the players on the left side of the court and visible blocking artefacts in the stands and foreground" src="https://gmelin.io/blog/ffoveated/foveated-frame.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;Frame from participant #4, source #2. The red dot marks the fixation point; the court around it stays crisp while the ranks and the foreground crowd dissolve into macroblocks.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The frame above is the whole thesis in one image. All the action is confined to the left side of the court, the viewer is locked onto it, and the stadium ranks — which occupy most of the pixels — have been quantized into mush that nobody looked at closely enough to notice. Sports content like this is close to the ideal case; the savings are content-dependent, ranging from 54% on the busiest clip to 71% on the most forgiving one.&lt;/p&gt;
&lt;p&gt;For context, the comparable numbers in the literature at the time were around 63% (Arndt and Antons) and around 42% (Illahi et al., for cloud-rendered game streaming). Direct comparison is genuinely difficult, since anchoring on the JND was novel and the field had no common benchmark — but the picture is consistent.&lt;/p&gt;
&lt;h2&gt;What the Eye-Tracker Told Me About the Participants&lt;/h2&gt;
&lt;p&gt;The most unexpectedly interesting data was not the bitrates but the gaze paths. Compare two participants watching the same clip — a cheetah pacing back and forth:&lt;/p&gt;
&lt;div class="figure-grid"&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="Scatterplot of fixation points tightly clustered in a horizontal band across the middle of the frame" src="https://gmelin.io/blog/ffoveated/gaze-expected.png"&gt;&lt;/p&gt;
&lt;figcaption&gt;Participant #1: fixations track the animal.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="Scatterplot of fixation points scattered widely across the entire frame including corners and edges" src="https://gmelin.io/blog/ffoveated/gaze-explorative.png"&gt;&lt;/p&gt;
&lt;figcaption&gt;Participant #6: fixations sprayed across the whole frame.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/div&gt;
&lt;p&gt;The first is exactly what you would hope for: a tight band of fixations following the region of interest. The second is a participant who worked out what the experiment was measuring and went hunting for artefacts in the periphery — large saccades terminating repeatedly in objectively uninteresting background regions. That is not natural viewing behaviour, and it systematically drags their reported threshold down.&lt;/p&gt;
&lt;p&gt;Which suggests something reusable beyond this study: eye-tracking data could serve as a &lt;em&gt;post-hoc filter&lt;/em&gt; on participants in quality assessment databases, separating people who watched from people who audited. As far as I know that had not been done in image or video quality assessment. I wrote that idea up separately, along with the observation that foveated coding gives you a strong prior on &lt;em&gt;where&lt;/em&gt; a visible distortion must be — which in principle lets you infer noticeability from gaze irregularities alone, without asking anyone to press a button at all. The dataset here was too small to train anything on it, but the hook is there.&lt;/p&gt;
&lt;h2&gt;Where It Goes&lt;/h2&gt;
&lt;p&gt;Bigger screens and higher resolutions should push the savings further, since a larger share of every frame lands in the periphery — with encoding time as the likely bottleneck. Peripheral vision is disproportionately sensitive to abrupt contrast changes, so blurring the coarsely quantized blocks would probably buy additional headroom. And the lab dependency is the real barrier to scale: approximate webcam-based eye-tracking would be plenty precise for something this spatially forgiving, and would put the whole approach within reach of crowdsourced studies and, more to the point, of ordinary video calls.&lt;/p&gt;
&lt;p&gt;The full thesis is available as a &lt;a href="/static/publications/msc_thesis_oliver_wiedemann.pdf"&gt;PDF&lt;/a&gt;. Parts of it were published as &lt;a href="/static/publications/wiedemann2020foveated.pdf"&gt;&lt;em&gt;Foveated Video Coding for Real Time Streaming Applications&lt;/em&gt;&lt;/a&gt; at QoMEX 2020, and the participant-filtering ideas as &lt;a href="/static/publications/wiedemann2020gaze.pdf"&gt;&lt;em&gt;Gaze Data for Quality Assessment of Foveated Video&lt;/em&gt;&lt;/a&gt; at the ET-MM workshop at ACM ETRA 2020.&lt;/p&gt;</content><category term="science"/><category term="computervision"/><category term="videocoding"/></entry><entry><title>Local NR-IQA using Convolutional Neural Networks</title><link href="https://gmelin.io/blog/localiqa/" rel="alternate"/><published>2018-05-05T00:00:00+02:00</published><updated>2026-09-01T17:00:29.146903+02:00</updated><author><name>Oliver Gmelin</name></author><id>tag:gmelin.io,2018-05-05:/blog/localiqa/</id><summary type="html">&lt;p&gt;Teaching a CNN to judge image quality on 64x64 patches.&lt;/p&gt;</summary><content type="html">&lt;p&gt;&lt;strong&gt;Teaching a CNN to judge image quality one 64x64 patch at a time.&lt;/strong&gt;&lt;br&gt;
Perceptual quality has almost always been studied as a property of a whole image. My BSc thesis asked what happens if you zoom in instead.&lt;/p&gt;
&lt;h2&gt;Error Is Not Quality&lt;/h2&gt;
&lt;p&gt;The instinctive way to measure how badly an image has been degraded is to compare it to the original and average the squared differences. Mean squared error, and the peak signal-to-noise ratio derived from it, are still everywhere in the compression literature, largely because they are simple and have a tidy relationship to signal energy.&lt;/p&gt;
&lt;p&gt;They are also close to useless as a model of what a person sees:&lt;/p&gt;
&lt;figure class="w-wide"&gt;
&lt;p&gt;&lt;img alt="Four grayscale photos of a finch: the original, a version with Gaussian noise, one with salt-and-pepper noise, and one heavily blurred" src="https://gmelin.io/blog/localiqa/equal-mse.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;Every distorted version here has an identical MSE of 480 relative to the original. They are not remotely equal in perceived quality.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;All three degraded birds sit at exactly the same distance from the original under MSE. Nobody would rank them equally. The metric is pixelwise and sign-independent, so it is blind to the structure that makes natural images legible — and, as the literature had already established, unreliable both across content and across distortion types.&lt;/p&gt;
&lt;p&gt;Image Quality Assessment (IQA) is the field that tries to do better: predict the score a human panel &lt;em&gt;would&lt;/em&gt; give, without convening one. The hardest and most useful flavour is &lt;strong&gt;no-reference&lt;/strong&gt; IQA, where there is no pristine original to compare against — you get one image and have to judge it blind. That is the setting for anything real: photos uploaded to a platform, frames arriving over a stream, a compressor deciding how hard to squeeze.&lt;/p&gt;
&lt;h2&gt;One Number Per Image Isn&amp;rsquo;t Enough&lt;/h2&gt;
&lt;p&gt;Where I thought the field had left something on the table was locality. Essentially every method, classical or learned, collapses an image into a single scalar Mean Opinion Score — regardless of whether the photo is uniformly sharp, or a razor-focused subject sitting in a mushy out-of-focus background.&lt;/p&gt;
&lt;p&gt;A global score cannot tell you &lt;em&gt;where&lt;/em&gt; an image is good. Some methods operated on patches internally, but were only ever evaluated globally, and — more importantly — were trained on the assumption that an image&amp;rsquo;s global MOS is a valid label for every patch sampled from it. That assumption is obviously false for exactly the images where locality matters most.&lt;/p&gt;
&lt;p&gt;So the thesis set out to do three things: build a dataset with genuinely local labels, train a no-reference predictor on it, and then check whether local predictions are still useful once you zoom back out to global tasks.&lt;/p&gt;
&lt;h2&gt;How Small Is a Quality Judgement?&lt;/h2&gt;
&lt;p&gt;First: what does &amp;ldquo;local&amp;rdquo; even mean? Two constraints pull against each other. A single pixel plainly carries no quality information, and small patches are hard for humans to assess at all. But a large patch starts smuggling in &lt;em&gt;content&lt;/em&gt; — and once a rater can see what the picture is of, their judgement gets contaminated by whether they like the subject and how they think it ought to be depicted.&lt;/p&gt;
&lt;figure class="w-wide"&gt;
&lt;p&gt;&lt;img alt="A dragonfly photo above four crops of increasing size, from an unrecognisable 32x32 square to a clearly readable 224x224 crop" src="https://gmelin.io/blog/localiqa/patch-sizes.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;Content perceptibility against patch size. At 32x32 there is nothing to go on; at 224x224 — the standard ImageNet input — you are looking at a photo of a dragonfly and judging it as such.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;I settled on &lt;strong&gt;64x64 pixels&lt;/strong&gt;, sampled from 1024x768 images: as small as possible while still assessable, and small enough to give away few content cues. It is an ad-hoc choice and I said so in the thesis — it is not defensible against every objection, but a decision had to be made.&lt;/p&gt;
&lt;h2&gt;Building KonPatch&lt;/h2&gt;
&lt;p&gt;No dataset of locally annotated patches existed, so the first contribution had to be one. &lt;strong&gt;KonPatch&lt;/strong&gt; is 32,000 individually annotated 64x64 patches: 500 source images drawn from &lt;a href="https://arxiv.org/abs/1803.08489"&gt;KonIQ-10k&lt;/a&gt;, 64 patches sampled at random locations from each. Those 500 images were then excluded from the rest of the database, so nothing downstream is ever evaluated on an image Patchnet partly trained on.&lt;/p&gt;
&lt;p&gt;Annotation was deliberately cheap: a binary lab judgement per patch — does this look like it came from a high-quality image, or not? One vote each. That is a compromise, and it introduces a hard decision boundary into a phenomenon that is clearly not binary.&lt;/p&gt;
&lt;p&gt;The trick that makes it workable is that every patch inherits a source image with a known global MOS. So the label becomes a continuous score: a patch flagged as high quality takes its parent image&amp;rsquo;s MOS, and a rejected patch takes zero. A binary study, scored continuously.&lt;/p&gt;
&lt;div class="figure-grid"&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="Grid of sixty small image patches with visible texture, edges and detail" src="https://gmelin.io/blog/localiqa/konpatch-hq.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;Flagged as indicative of a high-quality source.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="Grid of sixty small image patches that are mostly flat, blurred, dark or washed out" src="https://gmelin.io/blog/localiqa/konpatch-lq.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;Flagged as &lt;em&gt;not&lt;/em&gt; indicative.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/div&gt;
&lt;p&gt;The distinction is intuitive even stripped of context: crisp edges, plausible focus and clean colour on one side; blur, noise and washed-out flat regions on the other.&lt;/p&gt;
&lt;h2&gt;Patchnet&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Patchnet&lt;/strong&gt; is a 14-layer CNN: seven convolutional layers with three maxpooling stages between them, the final feature maps vectorised into a fully connected predictor of 1024, 16, 8 and finally 1 neuron. ReLU throughout, except the output, where a sigmoid bounds the prediction to [0, 1].&lt;/p&gt;
&lt;figure class="w-full"&gt;
&lt;p&gt;&lt;img alt="Diagram of the Patchnet architecture: a 64x64x3 input passing through conv and maxpool stages down to 8x8x16, then a fully connected head of 1024, 16, 8 and 1 neurons" src="https://gmelin.io/blog/localiqa/patchnet-arch.png"&gt;&lt;/p&gt;
&lt;figcaption&gt;Patchnet: 64x64x3 in, one quality score out.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Twenty percent of KonPatch was held out for testing, and the remaining 25,600 patches split five ways for cross-validation, with a separate training run from random initialisation per fold. Implementation was Keras on TensorFlow, Adam, batch size 512, trained on Nvidia K40s.&lt;/p&gt;
&lt;p&gt;Validation loss bottoms out and turns upward after roughly 30 epochs. Rather than adding dropout to a network this small, I used early stopping, keeping the best-performing checkpoint per fold. Those five models land within a narrow band of each other on the held-out test set: MSE between 0.059 and 0.071, SROCC between 0.649 and 0.678.&lt;/p&gt;
&lt;p&gt;Taken at face value, an MSE of 0.06 on labels spanning [0, 1] looks poor. It isn&amp;rsquo;t quite what it seems: because of how the scores were constructed, the label distribution is two well-separated clusters — rejected patches at zero, accepted patches up near their source MOS — with very little in between. The distance between the clusters is large compared to the error, so the model is not straddling the boundary the way that number suggests. The consistency of the early-stopped models transferring cleanly to unseen test data was the more meaningful signal.&lt;/p&gt;
&lt;h2&gt;From Patches to Maps&lt;/h2&gt;
&lt;p&gt;A patch scorer becomes something more interesting when you slide it. Run Patchnet across a whole image with stride δ and you get a quality &lt;em&gt;map&lt;/em&gt; — a spatial readout of predicted quality instead of a single number.&lt;/p&gt;
&lt;figure class="w-wide"&gt;
&lt;p&gt;&lt;img alt="Four panels: a photo of an impala in dry grass, Patchnet's blocky black-and-white quality map, a smoother wavelet sharpness map, and the grayscale image" src="https://gmelin.io/blog/localiqa/feature-maps.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;Input (a), Patchnet&amp;rsquo;s quality map (b), the FISH wavelet sharpness baseline (c), and plain luminance (d). Patchnet separates the sharp animal from the softly blurred surroundings.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2&gt;Do the Maps Match Humans?&lt;/h2&gt;
&lt;p&gt;Eyeballing a map and declaring it good is not evidence. But since no local quality benchmark existed, I had to build the yardstick too: 125 fresh images from KonIQ-10k, 30 participants each, asked to draw bounding boxes tightly around regions they considered high quality — or tick a box saying the image contained none.&lt;/p&gt;
&lt;figure class="w-full"&gt;
&lt;p&gt;&lt;img alt="Left: a butterfly on orange flowers overlaid with many overlapping red rectangles. Right: the resulting smooth grayscale heatmap, brightest over the butterfly" src="https://gmelin.io/blog/localiqa/crowd-study.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;Thirty participants&amp;rsquo; bounding boxes (left) become a ground truth quality map (right) after normalising by participant count and smoothing.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Each pixel&amp;rsquo;s score is simply the fraction of participants whose boxes covered it, with repeat boxes from the same person counted once so overlapping selections cannot manufacture peaks. Because it is normalised by the number of participants rather than within the image, the resulting map expresses &lt;em&gt;absolute&lt;/em&gt; quality, not just relative differences — an image where nobody found anything good stays dark everywhere. Rectangles are a crude segmentation primitive, so the raw maps get a Gaussian smoothing pass with σ at 10% of the image width.&lt;/p&gt;
&lt;p&gt;Scoring predictions against this is not a straightforward MSE, since the two datasets are labelled differently. Instead I binarised the ground truth at a threshold, swept the prediction threshold, and measured the area under the resulting ROC curve — repeated across the whole range of ground truth thresholds. Median AUC sits around &lt;strong&gt;0.9&lt;/strong&gt;, with the lower quartile around 0.8 across the entire threshold range. There are genuine failures in the tail — individual images where the model does worse than chance — but the central tendency is clear: Patchnet&amp;rsquo;s local predictions track where people actually see quality.&lt;/p&gt;
&lt;h2&gt;Zooming Back Out&lt;/h2&gt;
&lt;p&gt;If local scores are meaningful, they should also say something about global quality. I ran the sliding window at stride 4 across the 9,500 KonIQ-10k images that KonPatch never touched, producing maps at about 5.5% of input resolution.&lt;/p&gt;
&lt;p&gt;The crudest possible aggregation — take the mean of the map — already reaches SROCC 0.667 on KonIQ-10k, against 0.560 for FISH, the wavelet sharpness metric applied the same way. Neither is a strong correlation, which is expected when you flatten a whole map to its average, but the ordering is informative.&lt;/p&gt;
&lt;p&gt;Doing it properly meant training a second model on the maps: a headless DenseNet-169 fed three stacked spatially-small inputs — the Patchnet map, a FISH sharpness map, and a downscaled grayscale image — with the classification head replaced by a 1x1 convolution and global maxpooling, so the model accepts any input resolution. With rotation and flip augmentation the 5,700 training images became 64,400 feature maps at 82 distinct resolutions.&lt;/p&gt;
&lt;p&gt;That meta-model reaches &lt;strong&gt;SROCC 0.79 / PLCC 0.81&lt;/strong&gt; on KonIQ-10k, comfortably ahead of the classical no-reference metrics of the day — BRISQUE at 0.70/0.70, SSEQ at 0.59/0.61, BIQI at 0.54/0.61 — and ahead of the earlier patch-based CNNs, KangCNN (0.63/0.67) and BosICIP (0.65/0.67).&lt;/p&gt;
&lt;p&gt;It does not, however, catch DeepBIQ (0.90/0.92) or DeepRN (0.92/0.95). That gap is the most interesting result in the thesis, because it falls along a clean line. Everything below it, mine included, builds a global score by aggregating predictions over spatially small regions. Everything above it sees large regions or the entire image at once, transfer-learned from ImageNet classification models. The obvious reading: perceptual quality is a &lt;em&gt;mixture&lt;/em&gt; of local, technical properties — sharpness, noise, compression artefacts — and global, content-dependent ones to do with composition and aesthetics. No patch, however cleverly chosen, can see the second kind. Deliberately throwing away the big picture costs you exactly the part of the judgement that depends on it.&lt;/p&gt;
&lt;h2&gt;Application: Letting Quality Maps Steer JPEG&lt;/h2&gt;
&lt;p&gt;The last chapter puts the maps to work. My group had previously built a &lt;a href="/static/publications/hosu2016saliency.pdf"&gt;variable-quality JPEG coder&lt;/a&gt; that adjusts quantization per 8x8 block according to a saliency map — with the significant catch that the saliency maps came from eye-tracking or crowd studies, one per image. That does not scale.&lt;/p&gt;
&lt;p&gt;The idea here was to substitute Patchnet&amp;rsquo;s quality maps for those human-generated saliency maps, on the hypothesis that high-quality regions and salient regions largely coincide: in a well-shot photo, the subject is in focus and the background is not.&lt;/p&gt;
&lt;figure class="w-full"&gt;
&lt;p&gt;&lt;img alt="Six panels showing the VarJPEG pipeline applied to a tiger cub photo: original, raw quality map, pruned binary map, cluster centroids, blurred centroids, and final contours" src="https://gmelin.io/blog/localiqa/varjpeg.jpg"&gt;&lt;/p&gt;
&lt;figcaption&gt;VarJPEG: predict a map, prune to the strongest 10%, cluster to 8 centroids, blur, and modulate per-block quantization with the result.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The clustering step is not cosmetic. Storing a per-block quality map inside the file would eat the bit budget it is supposed to save, so the map is pruned to its strongest 10%, reduced to eight k-means centroids, and reconstructed at decode time by blurring those centroids back out — meaning the file carries eight coordinate pairs rather than a full map.&lt;/p&gt;
&lt;p&gt;To find out whether it actually helps, I ran another crowd study on the same 125 images. For each source, a binary search over the global quality parameter produced 100 compressed versions with evenly spaced bitrates, and participants dragged a slider to degrade a test image beside the unmodified reference until they first noticed a difference — the just noticeable difference, 15 votes per image, with a standard JPEG run as the control.&lt;/p&gt;
&lt;p&gt;It was a wash. No consistent bitrate win emerged, and at very low bitrates VarJPEG was systematically &lt;em&gt;worse&lt;/em&gt; — the fixed centroid overhead is a larger share of a smaller file. There was no clear trend against overhead share, source bitrate, or MOS either.&lt;/p&gt;
&lt;p&gt;The likely explanation is a good lesson about evaluation design. VarJPEG deliberately starves the background, so at equal average bitrate it puts &lt;em&gt;more&lt;/em&gt; artefacts there than plain JPEG does. The interface then made that worse: dragging the slider makes the first distortions appear as a flicker, and they appear precisely in the regions the algorithm sacrificed. Participants were reporting distortions honestly — the task just pointed their attention at the background, which is exactly where the method trades quality away. The measurement was aimed at the wrong part of the frame.&lt;/p&gt;
&lt;h2&gt;Looking Back&lt;/h2&gt;
&lt;p&gt;The core claim held up. Quality is meaningfully local, a CNN can learn it from small patches with genuinely local labels, and those local predictions carry real information even when the eventual task is a global score. The parts that did not pan out were the more instructive ones: the gap to whole-image models pointing at what patches structurally cannot see, and the VarJPEG study showing how easily a subjective experiment can measure the wrong thing.&lt;/p&gt;
&lt;p&gt;The full thesis — architectures, training curves and the complete correlation tables — is available as a &lt;a href="/static/publications/bsc_thesis_oliver_wiedemann.pdf"&gt;PDF&lt;/a&gt;. Parts of this work were published at QoMEX 2018 as &lt;a href="/static/publications/wiedemann2018disregarding.pdf"&gt;&lt;em&gt;Disregarding the Big Picture: Towards Local Image Quality Assessment&lt;/em&gt;&lt;/a&gt;.&lt;/p&gt;</content><category term="science"/><category term="computervision"/><category term="machinelearning"/></entry><entry><title>Graphcx</title><link href="https://gmelin.io/blog/graphcx/" rel="alternate"/><published>2017-05-08T00:00:00+02:00</published><updated>2026-09-08T19:08:38.462142+02:00</updated><author><name>Oliver Gmelin</name></author><id>tag:gmelin.io,2017-05-08:/blog/graphcx/</id><summary type="html">&lt;p&gt;A collection of graph analysis and drawing algorithms&lt;/p&gt;</summary><content type="html">&lt;p&gt;&lt;strong&gt;A collection of graph analysis and drawing algorithms&lt;/strong&gt;&lt;/p&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="tree82-galaxy" src="https://gmelin.io/blog/graphcx/tree82-galaxy.png"&gt;&lt;/p&gt;
&lt;figcaption&gt;Galaxy layout of &lt;a href="/static/graphs/tree82inc.txt"&gt;tree82inc&lt;/a&gt;.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2&gt;History&lt;/h2&gt;
&lt;p&gt;I began working on graphcx as a project on genetic algorithms at the University of
Limerick in 2016. The mandatory implementation was largely extended with a flexible
OpenGL based visualization module and bunch of interesting algorithms.&lt;/p&gt;
&lt;p&gt;The probably most interesting part is an implementation of the &lt;em&gt;Sampled Spectral
Distance Embedding&lt;/em&gt; algorithm, which was presented by Çivril, Magdon-Ismail
and Bocek-Rivele at the 14th International Symposium on Graph Drawing, Karlsruhe, 2006.
&lt;a href="http://dx.doi.org/10.1007/978-3-540-70904-6_5"&gt;http://dx.doi.org/10.1007/978-3-540-70904-6_5&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The euclidean distances between vertices in a SSDE layout approximate the underlying
graph-theoretical distances. An obvious application is to combine SSDE on a global scale
with a further local optimizer (e.g. a force-directed algorithm) to speed up the drawing
process and circumvent bad local energy minima.&lt;/p&gt;
&lt;h2&gt;About the Project&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Written in Java, depends on &lt;a href="http://la4j.org"&gt;la4j&lt;/a&gt; and &lt;a href="https://jogamp.org/jogl"&gt;jogl&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Build system / dependency management: &lt;a href="https://gradle.org/"&gt;gradle&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Source: &lt;a href="https://github.com/crxxn/graphcx"&gt;GitHub&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Functionality &amp;amp; Algorithms&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Genetic optimization by simulated annealing:&lt;ul&gt;
&lt;li&gt;Grid or circular layout&lt;/li&gt;
&lt;li&gt;\(\sum\limits_i ||e_i||_2\) or #EdgeIntersections as fitness criteria.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Galaxy: A straight-forward &lt;em&gt;force directed&lt;/em&gt; implementation.
  The name is due to the typical shape this algorithm produces.&lt;/li&gt;
&lt;li&gt;Fruchterman-Reingold algorithm.&lt;/li&gt;
&lt;li&gt;Dijkstra shortest-path.&lt;/li&gt;
&lt;li&gt;General graph handling, representation and conversion.&lt;/li&gt;
&lt;li&gt;OpenGL based visualization which nicely depicts optimization progress.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Todo&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Add rudimentary interface to avoid recompiling b/c of parameter changes.&lt;/li&gt;
&lt;li&gt;Port OpenGL to Vulcan just for fun and out of interest.&lt;/li&gt;
&lt;li&gt;Add more algorithms.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Sample Layouts:&lt;/h2&gt;
&lt;div class="figure-grid"&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="tree82-square-edgelength" src="https://gmelin.io/blog/graphcx/tree82-square-edgelength.png"&gt;&lt;/p&gt;
&lt;figcaption&gt;Simulated annealing grid layout of &lt;a href="/static/graphs/tree82inc.txt"&gt;tree82inc&lt;/a&gt;, using $\sum\limits_i ||e_i||_2$ as a fitness criterion.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="tree82-ssde" src="https://gmelin.io/blog/graphcx/tree82-ssde.png"&gt;&lt;/p&gt;
&lt;figcaption&gt;SSDE layout of &lt;a href="/static/graphs/tree82inc.txt"&gt;tree82inc&lt;/a&gt;.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="matrix1-square-edgelength" src="https://gmelin.io/blog/graphcx/matrix1-square-edgelength.png"&gt;&lt;/p&gt;
&lt;figcaption&gt;Grid layout of &lt;a href="/static/graphs/matrix1inc.txt"&gt;matrix1inc&lt;/a&gt;, using $\sum\limits_i ||e_i||_2$ as a fitness criterion.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure&gt;
&lt;p&gt;&lt;img alt="original-galaxy" src="https://gmelin.io/blog/graphcx/original-galaxy.png"&gt;&lt;/p&gt;
&lt;figcaption&gt;Galaxy layout of &lt;a href="/static/graphs/originalinc.txt"&gt;originalinc&lt;/a&gt;.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/div&gt;</content><category term="science"/><category term="programming"/><category term="graphs"/></entry><entry><title>Traceu</title><link href="https://gmelin.io/blog/traceu/" rel="alternate"/><published>2016-08-10T00:00:00+02:00</published><updated>2026-08-18T10:32:51.960104+02:00</updated><author><name>Oliver Gmelin</name></author><id>tag:gmelin.io,2016-08-10:/blog/traceu/</id><summary type="html">&lt;p&gt;A piecewise linear tracing utility for implicitly defined curves&lt;/p&gt;</summary><content type="html">&lt;p&gt;&lt;strong&gt;Piecewise-linear tracing utility for implicitly defined curves&lt;/strong&gt;&lt;/p&gt;
&lt;figure class="w-wide"&gt;
&lt;p&gt;&lt;img alt="traceu-plot1" src="https://gmelin.io/blog/traceu/traceu-plot1.png"&gt;&lt;/p&gt;
&lt;figcaption&gt;A trace of \(H^{-1}(0)\) for \(n=10\).&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2&gt;Example: Fixpoints&lt;/h2&gt;
&lt;p&gt;Let \( g: \mathbb{R}^n \rightarrow \mathbb{R}^n, g_k(x) = \exp\Big( \cos\Big( k \cdot \sum\limits_{i=1}^n x_i \Big)\Big) \) for \(k = 1,\ldots,n\)&lt;/p&gt;
&lt;p&gt;Considering the homotopy&lt;/p&gt;
&lt;p&gt;$$H:[0,1] \times \mathbb{R}^n \rightarrow \mathbb{R}^n, (\lambda,x) \mapsto x - \lambda g(x)$$&lt;/p&gt;
&lt;p&gt;it is obvious, that \(H(0,x) = x\) and \(H(1,x) = 0 \Leftrightarrow x = g(x)\)
One approach to find fixpoints of \(g\) is therefore to follow \(H^{-1}(0)\) starting at \(x=0\) until
\(\lambda = 1\) holds. &lt;strong&gt;traceu&lt;/strong&gt; provides piecewise linear approximations of such curves with upper
bounds for the error enabling further (e.g. newton-type) methods to be successfully applied
on the results. The picture above depicts a trace of \(H^{-1}(0)\) for \(n=10\).&lt;/p&gt;
&lt;h2&gt;About the Project&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Still beta, development is quite stale.&lt;/li&gt;
&lt;li&gt;Written in C.&lt;/li&gt;
&lt;li&gt;Quite unrestrictive regarding input requirements (e.g. no derivability needed).&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;Todo&lt;/h4&gt;
&lt;ul&gt;
&lt;li&gt;Multithreading.&lt;/li&gt;
&lt;li&gt;An input parser to get rid of recompiling for every problem.&lt;/li&gt;
&lt;li&gt;A GUI would be nice, something written in ncurses would do perfectly fine.&lt;/li&gt;
&lt;/ul&gt;</content><category term="science"/><category term="programming"/></entry></feed>