193 KiB
AVIF and AV1 implementation plan
Goal
Complete a production-quality, fully managed AV1 codec and its bounded AVIF/HEIF image container integration for ImageSharp. The finished work must decode and encode still images and bounded image sequences, preserve source precision, use ImageSharp memory ownership, and provide SIMD-first hot paths with one behaviorally identical scalar fallback.
This plan is the authoritative delivery checklist. A source file, unit test, build, self-roundtrip, or local implementation is not completion evidence by itself.
Takeover audit: 2026-09-05
The earlier checked boxes and measurements below are historical checkpoint reports, not accepted conclusions about the current encoder or complete decoder. The fresh production-path audit is still in progress. No benchmark has been run during this investigation, and the complete 738-file upstream diff has not yet received a line-by-line audit.
Reference and worktree evidence
- Live
git ls-remoteidentifies officialhttps://aomedia.googlesource.com/aommain asd565eec60f084421fa34fc0534b760c6452b6a6c. The export atD:\GitHub\ynse01\aom-d565eec6-sourcewas compared with that revision's official archive: 1,522 files present, 18 byte-identical, 1,504 differing only by CRLF versus LF, and zero remaining content differences. The archive is temporary and outside this repository. - Live ImageSharp main and local
upstream/mainboth resolve toadb982081a7e89a824f873f7f286f517e04f80dd. The starting HEAD isa7f0fca6b01d0d498862aba66e024942393c1614; the index is empty. The initial worktree contains 13 modified files and two untracked benchmark files. Preserve all existing work while correcting demonstrated defects. - The optimized native build is
D:\GitHub\ynse01\aom-d565eec6-build-x64-release: Ninja, MSVC x64, Release/O2 /Ob2 /DNDEBUG, runtime CPU dispatch, decoder, encoder, and high-bit-depth support. Generatedconfig/aom_config.henables SSE2, SSE4.1, AVX2, and AVX512. The cache's zero-valued HAVE entries do not describe the generated configuration. The version header reports 3.15.0 but does not independently identify a commit. - Commit
a7f0fca6balready contains temporary native integration:tests/ImageSharp.Benchmarks/Codecs/Heif/Native/aom_benchmark.c, itsCMakeLists.txt,LibaomBenchmarkEncoder.cs, and the referencing sequence benchmark. The current untracked decoder benchmark and adapter wrapper are also temporary integration. Do not stage or commit them. Removal from existing commits or deletion of local reference files requires a separate, concrete proposal; no history rewrite or deletion is authorized here. Existing generated conformance fixtures also require classification before any cleanup proposal.
Confirmed implementation deviations
Managed paths below are relative to the repository; reference paths are relative to the verified libaom export. Line numbers describe the inspected starting tree, before subsequent corrections.
| Class | Managed evidence | Official reference evidence | Finding |
|---|---|---|---|
| Missing functionality | Av1FrameEncoder.cs:372-408, under src/ImageSharp/Formats/Heif/Av1/Pipeline |
av1/encoder/encoder.c:641-646; av1/av1_cx_iface.c:287-288,1284-1286,1561-1562 |
Sequence setup unconditionally disables CDEF, restoration, and intra-edge filtering. These are not equivalent to the reference's configured tool decisions. |
| Architectural deviation | Av1FrameEncoder.cs:508-542,1551-1557 |
av1/encoder/encode_strategy.c:168-230,1664-1669 |
Every frame is error resilient, refreshes all slots, disables frame-end CDF publication, and resets probabilities. The reference selects retained primary-reference state. |
| Missing functionality | Av1IntraSuperblockEncoder.ModeDecision.cs:185-245; Av1IntraSuperblockEncoder.ReferenceModeDecision.cs:526-566 |
av1/encoder/partition_search.c:3320 onward; av1/encoder/rdopt.c:6196-6236 |
Inter frames retain a fixed 8x8 partition tree and search only LAST. Larger partitions and additional reference roles are not implemented by this path. |
| Architectural deviation | Av1IntraSuperblockEncoder.ModeDecision.cs:597-609,662-672,824-834 |
av1/encoder/rdopt.c:111-142,6186-6236; av1/encoder/intra_mode_search.c:1291-1344 |
Managed coding finishes intra search before inter evaluation. Reference inter search has its own ordered candidates, pruning state, bounds, and later intra evaluation. |
| Architectural deviation | Av1IntraSuperblockEncoder.ReferenceModeDecision.cs:576-617 |
av1/encoder/rdopt.c:111-142 |
Managed single-reference mode order is NEAREST, NEAR, GLOBAL, NEW. The reference default order is NEAREST, NEW, NEAR, GLOBAL across eligible references. |
| Architectural deviation | Av1IntraSuperblockEncoder.ReferenceModeDecision.cs:1278-1414 |
av1/encoder/mcomp.c; caller policy in av1/encoder/motion_search_facade.c |
The managed radius is an effort-shifted value capped by its border; each scale visits eight offsets once. This is a simplified search controller whose full reference-policy reconciliation remains open. |
| Missing functionality | Av1TransformBlockEncoder.cs:1047-1097 |
av1/encoder/encodemb.c:842-885 |
Lossy transform coding ends at fast quantization. Reference coding selects quantization with trellis policy and can optimize coefficients before reconstruction. Native primitive arithmetic alone does not establish encoder parity. |
| Architectural deviation | Av1IntraSuperblockEncoder.ModeDecision.cs:1581-1655 |
av1/encoder/intra_mode_search_utils.h:622-657; av1/encoder/intra_mode_search.c:467-491,1597-1619 |
The 8x8 SATD screen imports only the 1.5-best threshold. The reference also maintains ranked candidates and quantizer/neighbor-dependent pruning, and owns the surrounding mode/transform decisions. Contrary to an initial audit hypothesis, this revision's intra_model_rd does return raw SATD. No modeled-RD-versus-SATD numerical defect is established. |
| Missing functionality | Av1TransformBlockEncoder.cs:695-837 |
av1/common/reconintra.c:958-986,1204-1243 |
Encoder directional prediction bypasses edge preparation and upsampling. Flipping the sequence flag alone would make encoder reconstruction disagree with its emitted syntax. |
| Verification gap | Av1InverseTransformerFactory.cs:47-60,98-112; Av1InverseTransformTests.cs:402-532 |
Sparse inverse dispatch in av1/common/idct.c and av1/common/x86 |
The uncommitted decoder path specializes DC-only DCT; other lossy EOB values still use full transforms. Its new tests compare against the managed full transform, not an independent native oracle. Complete sparse dispatch and SIMD reconciliation remain open. |
Two reconstruction-input defects were established and corrected during this audit:
- Rectangular intra transforms require width plus height samples on each extended edge. The starting encoder
prepared twice the width above and twice the height to the left
(
Av1IntraSuperblockEncoder.ModeDecision.cs:1466-1534,2583-2677;Av1IntraSuperblockEncoder.ChromaModeDecision.cs:179-182,1153-1223). This could expose unprepared scratch samples to directional prediction. Referenceav1/common/reconintra.c:1149-1184,1451-1488,1817-1820copies the available adjacent edge and repeats its endpoint through the width-plus-height extent. Luma and chroma now reuse the existing edge-preparation method, and tiled candidate preparation follows the same extent without adding storage or an allocation. - Trial and final geometry discarded the enclosing partition, and all encoder directional edge-availability calls
passed
None(Av1IntraSuperblockEncoder.ModeDecision.cs:397-418,858-921,1434-1464,2557-2581in the starting tree). Referenceav1/common/reconintra.c:158-192,343-379selects different availability tables for mixed vertical partitions;av1/decoder/decodeframe.c:1392-1416distinguishes split child nodes from mixed-partition leaves. Both trial and final setup now retain the terminal partition in existing mode information, and luma, chroma, and tiled prediction consume it. Split children retain their own implicitNoneleaf state.
The isolated 8x8 Hadamard pruning hook has been removed from production mode selection. Its primitive and existing component tests remain uncommitted in the worktree. Removing the unsupported hook does not complete the remaining encoder controller or establish a quality or performance improvement.
Disabled tools, limited search, and different decision order can also change reconstructed samples without producing an invalid bitstream. They are separate from the reconstruction-input defects above.
Architecture and verification findings
Reference-edge correction checkpoint 578ec34d9, verified on 2026-09-05:
- Final Release .NET 11 build after removing the screening hook: zero errors and zero reported warnings on that incremental build. The preceding test compilation reported 1,009 existing warnings.
- Visual Studio VSTest 18.9, .NET 11 preview 7, serialized collections, one test thread, stop-on-failure:
251/251 cases passed in
Av1EncoderFrameTests,Av1IntraSuperblockEncoderTests, andHeifEncoderTests. The final report isD:\GitHub\ynse01\av1-takeover-20260905\no-screen-final.trx. RectangularIntraReferencesExtendTheLastAvailableSamplechecks explicit reference-edge samples at 8/10/12 bits for both rectangle orientations and available/unavailable extensions, including output sentinels.ProductionMixedPartitionsPreserveReconstructionOrderrequires actual mixed partitions and compares the live mapped encoder partition state and retained reconstruction with production decoding. Its two emitted 32x32 monochrome streams also match the optimized libaom decoder: 2,048 luma samples, maximum error 0, and zero samples exceeding one.- The twelve freshly regenerated two-frame color streams cover 8/10/12-bit 4:2:0, 4:2:2, and 4:4:4:
managed and optimized native decoding agree on all 21,348 Y/U/V samples, maximum error 0, zero exceeding one.
Per-frame and per-plane counts are in
D:\GitHub\ynse01\av1-takeover-20260905\decoder-comparison.json. - These are bounded reconstruction and same-bitstream decoder checks. They do not prove separately encoded output parity, complete encoder control flow, or decoder-wide conformance. No benchmark was run. Temporary launch scripts, native comparison output, logs, and reports remain local and are excluded from commits.
Sequence-construction ownership correction, verified subsequently on 2026-09-05:
-
At checkpoint
578ec34d9,Av1FrameEncoder.cs:1492-1560,1668-1713,1778-1824allocated common state and frame owners without unwinding partial construction.Av1EncoderPictureBuffer.cs:87-164andAv1SymbolEncoder.cs:270-274had the same problem inside their multi-owner constructors. A failed constructor never returns an instance to the caller'susingstatement. This is a demonstrated lifetime defect, independent of encoder quality. Referenceav1/encoder/encoder.c:1480-1499clears compressor state and invokesav1_remove_compressorif construction fails. -
The new production-factory regression failed before the correction: rejecting the second allocation left the first allocation unreturned. The first failure stopped VSTest as configured; evidence is
D:\GitHub\ynse01\av1-takeover-20260905\allocation-failure-before.trx. -
Common sequence state, both sample-width frame constructors, picture state, and symbol state now unwind completed owners at their own construction boundaries. Picture context views borrow its two owners (
Av1NeighborArrayUnit.cs:65-70,135-141), so failed picture construction releases those owners directly. Common cleanup avoids dispatching into derived frame disposal before derived construction begins. Nullable disposal checks are restricted to owners that can be absent after an allocation failure; no new owners, buffers, copies, or native dependencies were introduced. -
Six color/alpha cases at 8/10/12 bits reject every allocator request in turn and require exactly one return for every earlier successful request. All six pass. The final affected frame, superblock, and public encoder set passes 257/257 cases through serialized Release .NET 11 Visual Studio VSTest with stop-on-failure. The report is
D:\GitHub\ynse01\av1-takeover-20260905\ownership-final.trx. -
The final incremental Release .NET 11 build reports zero warnings and errors; Roslynk reports zero compiler errors. The twelve regenerated color sequences and two mixed-partition streams were compared again with the optimized native decoder: 23,396 samples, maximum error 0, zero samples exceeding one. This remains bounded same-stream decoder/reconstruction evidence, not separate-encoder parity. No benchmark was run.
-
PNG resolves options before converted metadata and sanitizes incompatible output combinations (
src/ImageSharp/Formats/Png/PngEncoderCore.cs:1632-1675). TIFF follows the same precedence and converts unsupported combinations (src/ImageSharp/Formats/Tiff/TiffEncoderCore.cs:111-175,371-457). HEIF's generic pixel input must retain that conversion contract; source pixel type is not an eligibility gate. -
JPEG's closed
JpegColorConverter<TOperator>owns traversal and calls semantic static operator arithmetic (src/ImageSharp/Formats/Jpeg/Components/ColorConverters/JpegColorConverter.Operator.cs:144-250). Shared prediction work must follow that family boundary and existing ownership APIs. -
Encoder block scratch already shares mode and inter storage by their non-overlapping lifetimes (
Av1EncoderBlockWorkspace.cs:31-103). The sequence constructor retains picture, coefficient, block, entropy, conversion, and frame state (Av1FrameEncoder.cs:1492-1646). Exact sizing, failure unwinding, alignment, reference parity, and allocation attribution still need the complete native comparison. -
Two decoders agreeing on one managed bitstream establishes only decoding agreement for that bitstream. The benchmark's photographic setup checks that agreement; it does not compare separately encoded outputs. Its native output uses ImageSharp color conversion, so it is not an independent RGB conversion oracle.
-
Historical separate-encoder Y/U/V maxima of 40/30/53 fail the required one-component-unit limit. Counts exceeding one were not supplied with those historical figures. Neither those figures nor the recorded 1,363.09/71.73 ms timing pair is a new measurement of a subsequently edited tree.
-
VSTest logs identify Visual Studio 18.9 x64 and
.NETCoreApp,Version=v11.0. Future runs must deduplicate child environment keys case-insensitively, disable collection parallelism, stop on failure, and run serially. Absence of an observed dialog is not evidence that no popup occurred.
Ordinary-intra skip investigation at checkpoint c778217a9:
- Managed
Av1IntraSuperblockEncoder.ModeDecision.cs:622-667,758-829,1291-1396scans retained empty transforms, computes their rate again, and can replace ordinary intra coefficient syntax with block skip.Av1TileWriter.cs:2628-2648explicitly describes this as a departure from current libaom. - Verified reference
av1/encoder/rdopt.c:3516-3578assigns ordinary intra rate includingskip_txfm_cost[skip_ctx][0]and setsskip_txfm = 0before separate IBC evaluation. Its intra-in-inter-frame path does the same atrdopt.c:5772-5801. Final coding enforces this atav1/encoder/partition_search.c:2122. Referenceav1/common/blockd.h:372-374counts IBC as inter for this decision. - This is an encoder-policy deviation, not an invalid-bitstream claim. Suppressing empty transform symbols
changes block-skip and coefficient probability adaptation for subsequent blocks even when current pixels match.
Managed
Av1TileWriter.cs:825-852,1100-1143consumes the selected flag in symbol order and publishes neighbors. - The regression
PreservesIntraNonSkipForAllZeroTransformsasserts the reference policy while retaining all zero-EOB and precomputed-versus-live syntax checks. It failed on the fixed-DC traversal before correction: mono intra returnedSkip = true(intra-skip-before.trxin the local takeover report directory). The first affected run then exposed a second path:Av1IntraSuperblockEncoder.cs:214-285independently marked zero-coefficient blocks skipped, so its exact byte comparison with the corrected live path failed. Both paths now preserve ordinary-intra non-skip, and their exact byte comparison is retained. These changes enforce the demonstrated reference contract; no pixel tolerance was changed. - Automatic approval review rejected a combined production patch and deletion of the old helper-specific
entropy test. A second review rejected their removal after verification as weakened coverage.
The unused helper and its test remain unchanged; no further removal was attempted.
With that helper and test still present, the corrected production, fixed-DC, and public encoder paths pass
258/258 serialized Release .NET 11 VSTest cases (
intra-skip-r2.trx). Decoded mode assertions cover ordinary intra syntax in key/inter frames at 8/10/12 bits and all three color subsampling formats; repeated-frame tests still require skipped inter blocks. The build has zero errors and 1,009 existing test-project warnings, with none in the changed files; Roslynk reports zero compiler errors. Optimized native decoding of freshly regenerated streams matches 23,396 samples: maximum error 0 and zero samples exceeding one. This is bounded same-stream evidence, not separate-encoder parity. Roslynk now finds only the old entropy test referencing the obsolete skip helper; no production caller remains. No benchmark was run, and the full controller, coefficient optimization, and separate-encoder gates remain open.
Coefficient optimization and evaluation-stage investigation:
av1/encoder/encodemb.c:208-224,474-563,831-887selects quantization and coefficient optimization from segment and evaluation policy before publishing coefficient context and reconstruction.av1/encoder/rdopt_utils.h:608-709distinguishes default, mode, and winner evaluation: transform pruning, default transform use, skip/DC prediction, distortion domain, coefficient optimization threshold, and transform-size search differ by stage. Changing stage invalidates cached RD results.av1/encoder/txb_rdopt.c:400-560consumes plane/block RD scaling, transform/EOB/context costs, quantized and original coefficients, dequantization and matrices. It can lower coefficients, move EOB, and select an empty transform; it updates coefficients, EOB, entropy context, and rate together. Importing this primitive without those callers and their state would leave the controller deviation unresolved.aom/aomcx.h:215-221,av1/av1_cx_iface.c:781-782,1401-1409, andav1/encoder/speed_features.c:2709-2776show usage-specific CPU settings and feature initialization. ImageSharp's 0-10 effort scale has not yet been reconciled with these policies. No new effort mapping is assumed.
Color-conversion boundary correction after checkpoint f7bd907d6, verified on 2026-09-05:
HeifEncoderCore.Sequence.cs:74-194resolved output sampling and preserved reversible YCgCo matrix metadata, including when default 4:2:0 or requested 4:2:2 was incompatible. The shared converter's established contract rejects that combination atHeifColorConversionParameters.cs:265-282. The public save regression failed on default sampling with that exact exception (matrix-fallback-before.trx).- PNG and TIFF resolve incompatible options through conversion at
PngEncoderCore.cs:1640-1654andTiffEncoderCore.cs:378-444. HEIF now extends its existing identity-matrix fallback to incompatible YCgCo-Re/Ro sampling, converts with BT.601, and writes the matching matrix metadata. Explicit 4:4:4 remains eligible for the existing reversible operator. This changes encoder option resolution, not decoder acceptance. - Seven public cases retain the original profile object and values, inspect decoded matrix/range metadata, and require byte-identical output to the same packed pixels explicitly encoded with the fallback matrix. The cases include the original identity fallback and YCgCo-Re/Ro with default, 4:2:0, and 4:2:2 sampling.
- A separate lifetime defect existed at
Av1FrameEncoder.cs:1404-1434: conversion parameters were resolved after renting row storage. The internal-factory regression confirmed that rejected conversion had already made one allocator request (conversion-allocation-before.trx). Resolution now precedes storage allocation; the regression requires both allocation and return logs to remain empty. No new guard or owner was added. - Final Release .NET 11 build: zero reported warnings and errors on the incremental build. Roslynk reports
zero compiler errors. Serialized Visual Studio VSTest with stop-on-failure passes 265/265 affected cases
(
conversion-final.trxin the local takeover report directory). The regenerated color and partition streams again match optimized native decoding on all 23,396 samples, maximum error 0 and zero exceeding one. No benchmark or separate-encoder parity comparison was run.
Color/output source coverage and remaining limits:
Av1YuvConverter.cs:24-135,175-274dispatches byte/high-bit-depth, complete/cropped/scaled output, and alpha through shared HEIF adapters.HeifPlanarColorConverter.cs:39-320,329-870was read through both traversals: conversion owns reusable row scratch, interpolates chroma before matrix conversion, and uses existingPixelOperations<TPixel>packing/unpacking or Rgb48/Rgba64 conversion.HeifColorConverter.Operator.cs:164-431uses closed semantic operators and descending SIMD widths. These architecture observations are not proof that every H.273 operator or pixel format is numerically correct. Independent color-conversion and complete SIMD coverage remain open.- The reference codec interface exposes native planes and strides (
aom/aom_image.h:284-292). The temporary comparison adapter supplies I420 using the managed RGB conversion. Its output cannot independently validate that RGB conversion, even when both codec decoders agree on native planes.
Required completion gates
Motion-controller investigation continued after correction checkpoint 578ec34d9:
-
Managed
Av1IntraSuperblockEncoder.ReferenceModeDecision.cs:1278-1493uses the same normalized squared-error plus complete mode/vector RD cost for integer and fractional candidates. The integer operators atAv1IntraSuperblockEncoder.Operator.cs:466-493,985-1014compute squared error, not SAD or centered variance. -
Reference
av1/encoder/mcomp.c:72-130,184-238,314-384,644-664separates full-pixel SAD cost, variance cost, SAD-per-bit scaling, error-per-bit scaling, and their motion limits. Its full-pixel dispatcher (mcomp.c:1768-1903) owns the configured diamond/hexagonal/pattern search and conditional mesh search, including downsampled-SAD fallback. -
Reference
av1/encoder/motion_search_facade.c:150-324,347-488derives the starting search step, considers configured start candidates, retains a second full-pixel candidate, prunes repeated dynamic-reference searches, and can refine and compare both candidates. Fractional search (mcomp.c:3266-3337) has configured precision, iteration count, repeated-position tracking, and a second-level check. The managed one-ring-per-scale controller does not implement that path. -
The next motion implementation must reconcile configuration, limits, start-candidate lifetime, search costs, full-pixel traversal, fractional traversal, and winner publication together. Substituting SAD or variance alone, increasing the radius, or adding isolated search points would not establish that contract. No benchmark or motion-search implementation change has been made from this follow-up investigation.
-
Reference good-quality speed policy is layered rather than a radius lookup. The defaults at
av1/encoder/speed_features.c:2353-2364select NSTEP, full eighth-pixel precision, two subpixel iterations, and eight-tap search. Good-quality overrides atspeed_features.c:1248-1250,1308-1312,1370-1409change iteration count, search range, full-pixel and fractional methods, second-candidate refinement, and mesh pruning. Resolution-dependent speed-six overrides atspeed_features.c:1029-1076also select block-size-dependent search and reference-candidate pruning. These source observations do not make existing managed effort values equivalent to native cpu-used values. -
Encoder intra-edge filtering requires the complete candidate prediction path. Reference
av1/common/reconintra.c:958-986,1204-1243,1512-1548derives neighbour-dependent strength, filters the corner and required edges, then upsamples before directional prediction. The decoder already performs these stages atsrc/ImageSharp/Formats/Heif/Av1/Prediction/Av1PredictionDecoder.cs:915-966. Encoder luma, chroma, and tiled candidates must consume the same filtered-reference contract before the sequence flag can be enabled. The decoder's chroma smooth-neighbour test atAv1PredictionDecoder.cs:1820-1837does not repeat the native inter-block check, butAv1TileReader.cs:2037-2042resets every inter block's UV mode to DC. That owning invariant prevents stale smooth modes; no redundant guard or numerical defect is justified here. -
Sparse inverse dispatch remains incomplete. Reference
av1/common/x86/highbd_inv_txfm_avx2.c:4088-4180derives separate horizontal and vertical nonzero extents from EOB and chooses low-one, low-eight, low-sixteen, or full DCT/ADST axis kernels as applicable. ManagedAv1InverseTransformerFactory.cs:47-60,98-112only distinguishes DC-only DCT from the full lossy transform. The DC arithmetic atAv1Inverse2dTransformer.cs:44-57retains separate axis scaling and rounding; agreement with the managed full path still does not independently establish all native sparse cases. -
Finish the full production-path and complete upstream-diff audit, including conversion, animation, ownership, filters, decoder SIMD, and independent validity of the claimed tests.
-
Reconcile frame configuration and encoder decision policy with the reference before isolated pruning changes.
-
Implement missing tools and complete reference, partition, motion, transform, coefficient, and winner decisions.
-
Compare separately encoded results from identical source samples with explicitly reconciled settings. Report maximum absolute error and counts exceeding one for every decoded output component and every frame. The acceptance limit is one component unit per sample; PSNR and average error cannot replace it.
-
Verify each final relevant edit with focused serialized Release .NET 11 Visual Studio VSTest and independent native production-output checks. Compilation and component tests do not close codec completeness.
-
Run equivalent end-to-end benchmarks only after the relevant source comparison justifies the next change. Retain output sizes, absolute times, per-sample errors, memory units, reference configuration, and limitations.
-
Inspect the staged diff before every verified checkpoint commit and exclude all temporary native integration, codec sources, binaries, build directories, and generated comparison artifacts. Do not push.
Source authority
- AV1 codec syntax, tables, fixed-point arithmetic, prediction, transforms, entropy behavior, filters, encoder decisions, and lifecycle behavior must be ported and checked only against the current
mainbranch of the official libaom checkout at the verifiedD:\GitHub\ynse01\aom-d565eec6-sourceexport. - Libaom is the sole external codec implementation source. Do not use HM, libheif, FFmpeg, GPAC, SVT-AV1, dav1d, libgav1, or any other codec implementation as an algorithm, arithmetic, output, or architecture reference.
- Existing ImageSharp and JPEG code is authoritative only for ImageSharp architecture, allocator ownership, SIMD dispatch, pixel conversion, and test API patterns. It is not an alternate AV1 algorithm source.
- Production code must not load, invoke, install, or fall back to a native codec.
- Existing independent container files may be used only as interoperability inputs. Native AV1 expected output must be generated by the current libaom
maincheckout, and no independent decoder output may substitute for it.
Reference checkout evidence refreshed on 2026-09-05:
- The current encoder source comparison uses the official libaom
mainrevisiond565eec60f084421fa34fc0534b760c6452b6a6c, exported atD:\GitHub\ynse01\aom-d565eec6-source. - The earlier generic reference decoder was
D:\GitHub\ynse01\aom-d565eec6-build-generic2\aomdec.exe; the fresh correction comparisons use the optimized x64 Release build recorded above. Its CMake cache identifies the current source export above, and it reports version 3.15.0. On 2026-09-05 it accepted both frames of each retained-reference sequence at efforts five, seven, eight, and nine. This establishes syntax acceptance for those four streams, not complete interpolation or codec conformance.
Status notation
- Recorded checkpoint: the associated dated report claims focused verification. Historical marks outside the takeover section have not been accepted by the fresh audit and do not establish current-tree completeness.
- [~] Locally implemented, checkpoint open: production source exists, but current-tree verification is missing or a known audit issue invalidates the checkpoint.
- Remaining: the production behavior is absent, incomplete, or has not reached its required implementation boundary.
Current source reconciliation
Reconciled with the worktree on 2026-09-05.
- [~] The bounded container reader, still-image path, sequence parser, AV1 decoder, color pipeline, presentation pipeline, and broad AV1 test suite exist locally.
- The inter-frame decoder has verified checkpoints through inter deblocking decisions and reference/mode deltas.
- [~] Loop filtering, CDEF, super-resolution, restoration, film grain, layered presentation, alpha composition, and color conversion exist locally. Shared-source cleanup changed the current tree, so final production-path verification is open.
- [~] AV1 writer primitives, forward transforms, symbol encoding, and tile-writing source are connected to the public encoder for bounded still-image and all-intra sequence AVIF color with optional auxiliary alpha output.
- [~] The public AV1 encoder has local single-image, grid, lossless sequence, and LAST_FRAME lossy sequence paths. Current-tree verification remains open. Additional references, compound prediction, remaining inter tools, orientation handling, and default format registration remain open.
- Patented codec production code, registrations, tests, benchmarks, fixtures, reference outputs, and notices were manually deleted and committed by
78a74d448. - Remaining task-created HM, HEVC, libheif, GPAC, Nokia, FFmpeg, Pillow HEIF, libavif-build, and libjpeg-build directories were traced to their creation commands in the recovered Codex session history and deleted on 2026-08-31. The user-provided repositories and all libaom-only source, build, and reference data were left untouched.
- The PNG metadata-suppression fix and three HEIF/AV1 diagnostic-save call-site corrections passed the exact 34 net11.0 ARM CI cases and were committed with the single-reference checkpoint as
54bb6cbe59bd113058854a3ee31448cf61f462ca. They are infrastructure evidence, not decoder or encoder completion evidence. - The complete decoder and encoder release matrix is not complete.
Immediate execution queue
Historical interpolation-search reports from before the takeover follow. Their timings and test counts apply only to the trees identified by those reports. The takeover audit and required completion gates above determine current work.
-
[~] The earlier elementary-stream decoder comparison included parsing, reconstruction, output allocation, RGB conversion, and disposal for both ImageSharp and optimized current-main libaom. Eight-bit Kodak and ten-bit Cosmos inputs match every
Rgb48sample before timing. Initial warmed means are 15.130 versus 5.423 ms and 27.655 versus 10.246 ms respectively. These expose an open decoder gap; they are not container-load measurements. Reproduction and limitations are recorded in the benchmark README. -
[~] Uncommitted shared inverse reconstruction uses EOB to select a DC-only DCT path, preserving both axis roundings, rectangular normalization, input clamps, and final clipping through the existing semantic output operators. Vector512 output was added to that operator contract. Every transform size, signed boundary, padded separate/in-place destination, and supported sample precision is checked against the full transform. The post-change photographic encoder payload remained byte-identical; decoder RGB output remained exact against libaom. Short decoder measurements do not yet establish a statistically significant improvement.
-
The isolated fixed-8x8 Hadamard screening hook was removed in correction checkpoint
578ec34d9. The following is a historical experiment, not an active implementation or an accepted improvement. Existing tensor and transpose APIs vectorize the operation over frame-reused scratch; no per-candidate allocation or custom hardware-width operator was introduced. A five-warmup, ten-measurement repeat records 1,363.09 ms versus native 71.73 ms, about 34% faster than the initial managed baseline but still approximately 19x behind native. Output increased slightly to 11.577 KiB and aggregate native-plane PSNR declined from 37.124 to 37.069 dB. This tradeoff does not close the performance/compression gate. Larger-block screening, top-ranked pruning, transform bounds, winner refinement, filter decisions, and allocation attribution remain open. Evidence:artifacts/BenchmarkDotNet/av1-screen-verified-short-20260905/20260905-140358. -
[~] The earlier screening/DC/native-profile/moving-color subset passed 25 cases in each of three separately configured VSTest hardware tiers. Normal-path verification passes 84 screening/DC/superblock cases and 165 frame/public encoder cases. The photographic benchmark now additionally requires exact ImageSharp/libaom RGB agreement across all three dependent frames before timing. No expected image or golden output was changed. Release test/benchmark builds have zero errors and their existing 1,009/39 warning baselines. Evidence and the corrected process-local VSTest environment handling are recorded in
tests/ImageSharp.Benchmarks/Codecs/Heif/README.mdandartifacts/TestResults/av1-screen-20260905. -
[~] The earlier reference benchmark measured RGB-to-OBU boundaries, including pixel conversion for every frame on both sides. A benchmark-only C adapter links the optimized current-main libaom build in-process and borrows the existing converted plane storage; production remains fully managed. Setup and file I/O are excluded equally. Native quantizer bounds are fixed to the managed base index, with independent speed settings and explicit size/quality reporting. Reproduction, native build provenance, lifetime documentation, and exact output hashes are in
tests/ImageSharp.Benchmarks/Codecs/Heif/README.md. -
The corrected benchmark exposes a substantial remaining performance and compression gap. For three photographic 256x256 frames, ImageSharp effort seven takes 2,053.31 ms and writes 11.54 KiB at 37.124 dB aggregate native YUV PSNR; current-main libaom cpu-used six takes 70.67 ms and writes 8.27 KiB at 38.942 dB. Both include RGB conversion and both measured outputs decode to all three complete frames. This is fixed-base-quantizer evidence, not equal-quality evidence. The managed path records 9.38 MiB of managed allocations per operation, which still requires attribution; native memory is not measured by that counter. The Short-run evidence is
artifacts/BenchmarkDotNet/av1-sequence-rgb-fixed-q-short-20260905/20260905-131149. Power-plan and CPU-query warnings remain documented. The new benchmark files build in net11.0 Release and have no Roslyn compiler/analyzer diagnostics. Do not close the interpolation-performance gate or advance to additional reference tools until the gap is addressed. -
The requested in-progress tree was committed as
433afd1a9before further encoder work. That commit is a checkpoint, not a claim of completed interpolation or codec delivery. -
The subsequent partition/interpolation/lossless correction was committed as
7cf7fc4after the 225-case affected encoder run and exact current-main native comparison. -
Odd-sized color sequence verification exposed and corrected two further production defects. Empty inter luma transforms now retain the inferred DCT type before chroma inherits it; normalizing only during writing was too late. Region-major coefficient writing now rounds chroma end coordinates in 4x4 units, preserving the shared minimum chroma transform on sub-8x8 partitions instead of truncating it away. Both rules match current libaom's transform-type inference and
av1_write_intra_coeffs_mbregion bounds. The corrections add no allocation or sample copy. -
All 12 moving-color sequence cases pass with SIMD enabled and disabled: 8/10/12-bit 4:2:0 at effort eight, and all three bit depths across 4:2:0/4:2:2/4:4:4 at effort nine. The 23x19 sources require actual inter motion and subsampled chroma phases. Current-main libaom decodes all 24 emitted frames with exact native Y/U/V equality. Evidence:
artifacts/TestResults/av1-interpolation-color-sequence-20260905/color-sequence-r3.trx,color-sequence-scalar-r3.trx, and the adjacent raw plane outputs undertests/Images/ActualOutput/Heif/Av1/SequenceEncoderPreservesNativeColorPlanesWithSubpixelMotion. The affected public/frame/superblock set passes 237 cases; the expanded encoder/entropy set passes 2,524, with no failures or skips. Evidence:artifacts/TestResults/av1-interpolation-partition-broad-20260905/color-broad-r3.trxandcolor-encoder-entropy-r3.trx. The final Release build and Roslyn compiler/analyzer passes have no errors. This closes the identified subsampled inter syntax/reconstruction gaps, not the remaining end-to-end performance and complete codec release matrix. -
Production interpolation verification now forces Smooth and Sharp at effort eight, both dual-filter axis orders at effort nine, and native 10/12-bit two-axis half-sample motion. The tests assert actual retained filter symbols, vectors, exact reconstruction, and no allocator rent during tile coding. The sequence fixture now uses the retained-reference decode contract; the still-image buffer-transfer API intentionally releases the reference map and cannot decode dependent samples in succession.
-
This verification exposed a live partition traversal defect: after search changed an earlier node's child count, a later node could consume an unrelated entry from the initial flat 8x8 skeleton. Unsearched intra partitions now derive their default from the current block size, preserving the geometry-driven traversal used by current libaom. Wide and tall lossless regressions cover clipped parents and superblock boundaries at efforts nine and ten.
-
The affected frame encoder, intra-superblock encoder, and public HEIF encoder set passes all 225 cases on the corrected tree with zero failures or skips. Evidence:
artifacts/TestResults/av1-interpolation-partition-broad-20260905/partition-broad-r3.trx. This is the affected encoder surface, not the full codec release matrix. -
Lossless tiled luma and chroma search no longer evaluate unsignaled angle deltas on 4x8/8x4 coding blocks; tiled chroma also respects the ordinary effort limits. Partition trials publish lossless chroma coefficient contexts at 4x4 transform granularity. The extended native RGB-plane regression exposed the chroma angle defect as a real lossless mismatch reproduced by libaom, and passes after the correction without changing expected samples.
-
All 20 focused partition/interpolation/lossless cases pass with hardware intrinsics enabled and all 20 pass with them disabled through serialized net11.0 Release Visual Studio VSTest with stop-on-failure. Current-main libaom
d565eec60f084421fa34fc0534b760c6452b6a6cdecodes all 20 emitted streams (28 frames) with exact native-plane equality against the retained reconstruction or lossless source planes. Evidence:artifacts/TestResults/av1-interpolation-partition-interpolation-20260905/partition-interpolation-r3.trx,partition-interpolation-scalar-r3.trx, andartifacts/av1-partition-interpolation-net11-20260905-r3.log. Roslyn compiler and analyzer passes report no errors or changed-file warnings. Subsampled inter-plane coverage, broader current-tree verification, and end-to-end performance remain open. -
The 8x8 single-candidate SAD, four-candidate SAD, and variance paths now share closed-generic traversal in
Av1ResidualBuilder. Its existing byte/ushort residual operators own scalar and SIMD arithmetic; the superblock operators no longer duplicate these row loops or hardware dispatch. The four-candidate path retains one source load/conversion per row, exact eight-sample loads preserve final-row bounds, and both variance moments retain native precision until the existing normalization boundary. Six known-result cases cover 8/10/12-bit signed extrema, distinct source/prediction rows and strides, unaligned starts, exact final-row lengths, and four-candidate output order. Together with nine existing intra-block-copy cases and the extended zero-allocation test, all 16 pass with hardware intrinsics enabled and all 16 pass with them disabled. All 58 public HEIF encoder cases also pass. Verification used serialized net11.0 Release Visual Studio VSTest with stop-on-failure; evidence isartifacts/TestResults/av1-search-metrics-20260905/search-metrics-r2.trx,search-metrics-scalar-r2.trx, andsearch-metrics-encoder-r2.trxin that directory. The exact Release build has zero errors and the existing 1,009 warnings; Roslyn reports no diagnostics in the changed files. This verifies the search-metric refactor, not the remaining interpolation conformance or end-to-end performance work. -
[~] Effort eight searches the common regular, smooth, and sharp interpolation families; efforts nine and ten enable independent vertical/horizontal filter selection. Lower efforts retain the fixed regular-filter path.
-
[~] Filter ranking follows the curve-fit prediction-error model in the official
d565eec60f084421fa34fc0534b760c6452b6a6csource, before full transform search. The existing rate-distortion type owns the static model tables and paired portable SIMD cubic evaluation. Visible-plane SSE uses the existing SIMD residual reduction, including native-bit-depth normalization and cropped edges. -
[~] Candidate and retained prediction views alternate within the existing inter workspace. At most one normalization copy per active plane retains the chosen predictor for transform search; no new pixel owner, coefficient owner, per-block rent, or reconstructed-frame copy is introduced. Zero-phase axes retain only the cheapest signaled filter instead of repeating equivalent prediction trials.
-
[~] The selected filters, skip flags, segment, and primary reference fit the original seven-byte block-mode record and eight-byte macroblock record. Filter costing and writing use the live tile CDFs. Encoder and decoder share the context-combination mapping, and encoder search and writing share filter-symbol eligibility.
-
[~] The earlier 67-case net11.0 Release set covered model curve samples and skip decisions, 8/10/12-bit quantizer normalization, exact entropy bytes and live adaptation, packed-field independence, tile-boundary contexts, syntax eligibility, and sequence-header signaling. Its 12 model cases also passed with hardware intrinsics disabled. The four flat retained-reference sequences were accepted by current libaom but did not force non-regular filters. The production, allocation, and exact native-plane evidence above extends that coverage; subsampled inter planes and end-to-end timing remain required before closing this checkpoint.
-
Focused runtime verification exposed an invalid low-effort inter-frame header:
force_integer_mvwas set while screen-content tools were disabled, making the writer omit the high-precision flag that a conforming reader expects. The frame encoder now retains the inferred false flag and controls integer-only search through the existing effort boundary. All four retained-reference cases assert the parsed precision, quantizer, and filter fields and decode both frames; current libaom accepts the same saved streams. -
The preceding interpolation checkpoint's affected net11.0 Release run passed 2,494 cases with zero failures or skips through one serialized Visual Studio VSTest process with stop-on-failure enabled. It combined 2,269 entropy cases, 167 frame/superblock/transform/picture-storage cases, and 58 public HEIF encoder cases. That build had zero errors and the existing 1,009 test-project warnings, with none in the changed files. Evidence:
artifacts/TestResults/av1-interpolation-20260905/verified-encoder-entropy-r10.trxandartifacts/av1-interpolation-net11-release-20260905-r10.log. This predates the search-metric refactor above and is not the complete current-tree decoder/encoder release matrix. -
Allocation regressions now reflect the implemented lifetimes instead of the former layout: all seven trailing tile-state integers are proved contiguous within the picture's second allocator owner; coefficient level, context, and output owners are proved allocated at construction, reused across costing, writing, finalization, and frame resets, and returned exactly once. The two-owner picture and three-owner symbol-encoder limits are retained.
-
Grid regressions use valid 4:2:2/4:2:0 syntax and the inferred monochrome subsampling flags used by production configuration. Odd-dimension and undersized-cell failures assert their grid-specific messages, so malformed AV1 fixtures or an earlier configuration mismatch cannot satisfy those tests. Smaller right/bottom color and alpha cells, separate primary roots, lossless sequences, metadata, and public precision/sampling cases pass in the 58-case set above.
Work must proceed in this order. Do not skip to a later item while an earlier checkpoint is open.
1. Finish and verify the AV1-only cleanup
- Remove production types, registrations, constants, parser branches, properties, tests, benchmarks, fixtures, reference outputs, notices, and documentation for removed codec work.
- Remove downloaded non-libaom reference source, tools, generated outputs, and local installations.
- Retain the official current-main libaom checkout and libaom-only build artifacts required for AV1 verification.
- Retain user-supplied AV1 fixtures and their recorded expected outputs.
- Audit production source, tests, benchmarks, assets, project files, notices, and documentation for stale removed-code references.
- The cleanup and cICP tree built in Release for net10.0 and net11.0 with restore disabled, build servers disabled, and one MSBuild node.
- The exact 34 net11.0 ARM CI failures pass after the cICP correction, and the subsequent single-reference checkpoint set passes on net10.0 and net11.0.
- Roslynk, scoped StyleCop, whitespace, and
git diff --checkaccepted the cleanup and cICP checkpoint. - The cleanup and cICP evidence was recorded and committed with the single-reference checkpoint.
Historical cleanup evidence from 2026-08-30, retained with its limitation:
- Release source builds passed for net10.0 and net11.0 with zero warnings and zero errors. Both builds used
--no-restore,--disable-build-servers, and one MSBuild node. - The focused net10.0 HEIF decoder, encoder, metadata, sequence-parser, and AV1 reconstruction set passed 221 of 221 tests with zero failures and zero skips. It did not execute the net11.0 diagnostic-save path that later failed in CI.
- The Roslyn compiler and configured StyleCop analyzers accepted the changed production source. Roslynk's
open_solutionentry point was attempted separately but failed before returning a solution handle, so no Roslynk result is claimed. - The tracked-source text and filename audit found no removed-code references outside the unchanged repository and shared-infrastructure
.gitattributespatterns. A later history reconstruction found ignored task-created reference directories that this audit missed; those directories were deleted on 2026-08-31. git diff --checkpassed and neither.gitattributesfile changed.
Current cICP failure correction evidence from 2026-08-31:
- The failure was not decoded HEIF metadata.
PngEncoderCore.WriteCicpChunkignoredPngChunkFilter.ExcludeAll, so diagnostic PNG saves attempted to write a non-identity source matrix that PNG cannot represent. PngEncoderCorenow honors the existingSkipMetadatacontract for cICP, and the three affected HEIF/AV1 diagnostic saves explicitly usePngEncoder { SkipMetadata = true }. Actual comparisons and decoded-image metadata assertions remain unchanged.- The direct embedded-ICC case and every row of the 12-case profile matrix passed: 13 of 13 net11.0 Release cases.
- The exact 34 cases reported by CI passed: 34 of 34 net11.0 Release cases, with zero failures and zero skips.
- Roslynk reported zero compiler errors after the fix, and
git diff --checkpassed. - The remaining four branch-introduced
GC.AllocateUninitializedArraycalls are removed from AV1 configuration, pixel-information, XMP, and Exif ownership boundaries. Each retained value still receives exactly one array and one copy because its source span belongs to pooled storage; no second materialization was introduced. The exact net11 Release rebuild remains at 1,005 warnings and zero errors, 166 focused configuration and metadata cases pass, and all 9,181 HEIF tests pass through direct VSTest.
Recovered task-history evidence from 2026-08-31:
- The primary session beginning on 2026-08-24 was reopened from task ID
01a03239-831b-7831-84e7-7f6947279ccb: 96,777 records, 295 turn contexts, 211 compactions, 190 user messages, 1,920 assistant messages, and 13,671 tool calls. - The continuation beginning on 2026-08-27 was reopened from task ID
01a04314-f1c6-7133-b1bc-5c74a94dd714: 61,129 records at the audit point, 166 turn contexts, 96 compactions, 223 user messages, 1,113 assistant messages, and 8,942 tool calls. - The restored first session records the user selecting official AOM/libaom as the AV1 source after the ImageSharp discussion was inspected. It does not authorize another codec implementation as an AV1 source and does not authorize importing a patented codec.
- The restored tool calls identify the exact creation commands for the non-libaom source, tool, and output directories removed on 2026-08-31. No directory was selected for deletion from its name alone.
- The recovered Git sequence establishes that
78a74d448removed the patented codec implementation and92fa7a8camerged the later upstream ImageSharp changes. The current branch and worktree, not an older summary, remain authoritative.
2. Correct the single-reference inter-frame checkpoint
The checkpoint is complete through c4b4e4e0386328dea574a884b6fa36c360ad5a9b. It replaces frame-sized palette maps with fixed decoder-session scratch, reconstructs each superblock before reusing that scratch, and passes the ownership, documentation, full AV1 test, and Release source-build gates on both target frameworks.
- Reconcile interpolation-filter syntax in
Av1TileReaderwith current libaommain.- Current libaom
av1_is_interp_neededcallsis_nontrans_global_motion, whose loop rejects onlyTRANSLATION. Identity GLOBALMV therefore omits switchable-filter symbols. - Current
Av1TileReaderuses the same non-Translation classification. The existing Identity test leaves sentinel filter symbols unread, while the Translation test consumes them. - No production change is required. The focused test describes only the syntax behavior it proves.
- Current libaom
- Reconcile both spatial single-reference extension loops in
Av1ReferenceMotionVectorswith current libaommain.- Current libaom
setup_ref_mv_liststops both loops atMAX_MV_REF_CANDIDATES, which is two.MAX_REF_MV_STACK_SIZE, which is eight, is the stack capacity used by the earlier direct and temporal candidate collection; it is not the stop condition for these two extension loops. - Current
Av1ReferenceMotionVectorsuses the same two-entry stop condition and retains an eight-entry stack for earlier candidates and DRL selection. - No production change is required. This remains spatial single-reference extension, not temporal extension.
- Current libaom
- Establish and enforce the contiguous frame-plane invariant used by
Av1FrameBufferand inter reconstruction.- One ImageSharp allocator owner now contains the aligned Y, U, and V storage, matching libaom's frame-buffer ownership while non-owning
Buffer2Dviews preserve ImageSharp's row API. Coded dimensions are aligned to eight samples, the luma stride is aligned to 32 samples, and chroma strides and heights are derived from that luma layout exactly once. A 4K eight-bit 4:2:0 frame owner occupies about 17.3 MiB. - The single owner removes the previous three-rent constructor and its allocation-cleanup
try/catch.Av1FrameBufferrejects external geometry whose complete aligned frame reaches the contiguousint.MaxValueboundary before allocation, making every directDangerousGetSingleSpancall an enforced owner invariant. ConstructorRequestsContiguousPaddedPlanesproves that a frame larger than the allocator's group capacity remains one group.ConstructorUsesOneFrameOwnerForAllPaddedPlanesproves exact one-rent Y/U/V ownership and exactly-once return.ConstructorRejectsPaddedPlaneThatCannotBeContiguousproves that an unrepresentable frame is rejected before allocation, and the high-bit-depth stride regression proves the 608-sample libaom layout for a three-pixel coded row.- The complete HEIF/AV1 namespace passes 8,808 of 8,808 direct net11 VSTest cases in Release after the physical layout change. The production path performs no plane copy and no per-block, per-row, or per-scanline allocation.
- One ImageSharp allocator owner now contains the aligned Y, U, and V storage, matching libaom's frame-buffer ownership while non-owning
- Prove the real
Av1BlockDecoder.DecodeBlockinter-reconstruction branch.- Decode the progressive dependent-frame fixture through the complete public production path.
- Compare the final frame's native Y, Cb, and Cr planes exactly with current-main libaom output.
- Compare the final presented image through the established ImageSharp reference-image comparison API.
- Do not substitute an internal helper test, fake tile reader, non-zero assertion, custom pixel loop, or tolerant comparison.
- Prove motion-field ownership and lifetime after the current reconstruction-timing change.
- Track initialization, retained-slot aliases, failure unwinding, presentation ownership, decoder-result ownership, and final disposal.
- Every allocator-owned object must be returned exactly once.
- Correct stale documentation for the current worktree.
- Av1InterFrameModeInfoTests must describe the behavior it actually proves.
- Do not claim production reconstruction, constrained allocation, ownership, or reference-stack coverage unless the test executes that contract.
Checkpoint gate:
- Default Identity-omission and Translation-consumption GLOBALMV syntax cases pass in the focused current-tree run.
- Two-entry spatial single-reference extension passes; current-main source inspection confirms the separate eight-entry overall stack capacity and DRL access.
- The exact dependent-frame native-plane comparison passes.
- The established exact presentation comparison passes.
- Normal, AVX-512-disabled, AVX-disabled, and scalar FeatureTestRunner configurations pass where supported.
- Constrained allocation preserves the enforced single-group plane invariant without copying or per-block allocation.
- Motion-field allocation tracking is balanced across success and failure on net10.0 and net11.0.
- Release source builds pass for net10.0 and net11.0 with zero warnings and zero errors.
- The complete AV1 namespace passes 8,732 of 8,732 tests on net10.0 and net11.0 with zero failures or skips.
- Roslynk reports zero compiler errors; scoped analyzer inspection reports no diagnostics introduced by the current changes;
git diff --checkpasses. - The completed checkpoint was committed as
54bb6cbe59bd113058854a3ee31448cf61f462cawith author and committerJames Jackson-South <james_south@hotmail.com>. - The palette-memory follow-up was committed as
c4b4e4e0386328dea574a884b6fa36c360ad5a9bwith author and committerJames Jackson-South <james_south@hotmail.com>.
Verified single-reference checkpoint evidence on 2026-08-31:
- The current-main
aomdecwas rebuilt directly fromD:\GitHub\AOMediaCodec\aomand identified itself as3.15.0-13-g441c439b99. - Decoding the 72-byte progressive payload with
--all-layers, one thread, and row multithreading disabled produced 2,178 YUV444 color samples. All samples in both layers match the first three planes of the stored YUV444-alpha reference exactly. DecodeProgressiveSingleMatchesReferenceexecutes the production decoder through FeatureTestRunner and compares the complete presentedRgba32image withCompareToReferenceOutput(ImageComparer.Exact, provider). The redundant manual alpha loop was removed.DecodeProgressiveSingleWithConstrainedAllocatorexecutes the same production reconstruction with a 1,024-byte allocator group capacity and verifies that every allocation is returned exactly once.MotionFieldsFollowAliasesAndPresentationOwnership,MotionFieldAllocationFailureUnwindsTileReaderOwnership,DecodeProgressiveSingleTracksMotionFieldOwnership, and the reference-store replacement, reset, and transfer tests cover initialization, aliases, presentation ownership, decoder-result ownership, failure unwinding, repeated disposal, and exactly-once final returns in the current worktree.- The current worktree passes the four-case palette set, seven-case ownership set, and 29-case syntax, plane, and production reconstruction set on both target frameworks. The complete AV1 namespace passes 8,732 of 8,732 tests on net10.0 and net11.0 with zero failures or skips.
- Release source builds passed for net10.0 and net11.0 with zero warnings and zero errors.
- Roslynk reported zero compiler errors. The scoped changed-file analyzer inspection reported no StyleCop diagnostics attributable to this checkpoint; its only remaining match is the pre-existing xUnit cancellation warning in an unrelated
HeifDecoderTestsmethod. git diff --checkpassed, and neither.gitattributesfile changed.
Exact verification commands, run directly in the foreground from D:\GitHub\ynse01\ImageSharp:
$env:MSBUILDUSESERVER = '0'
$env:DOTNET_CLI_USE_MSBUILD_SERVER = '0'
$env:DOTNET_CLI_HOME = 'D:\GitHub\ynse01\ImageSharp\.dotnet'
$env:DOTNET_SKIP_FIRST_TIME_EXPERIENCE = '1'
$env:DOTNET_CLI_TELEMETRY_OPTOUT = '1'
$env:DOTNET_DbgEnableMiniDump = '0'
$env:COMPlus_DbgEnableMiniDump = '0'
$env:DOTNET_EnableCrashReport = '0'
$env:COMPlus_EnableCrashReport = '0'
$heifCheckpointFilter = 'FullyQualifiedName~Av1InterFrameModeInfoTests.ReadInterFrameModeInfoReadsInterpolationFilters|FullyQualifiedName~Av1InterFrameModeInfoTests.IdentityGlobalMotionOmitsInterpolationFilters|FullyQualifiedName~Av1ReferenceMotionVectorsTests.BuildReversesOppositeDirectionExtensionCandidate|FullyQualifiedName~Av1FrameBufferTests|FullyQualifiedName~Av1ReferenceFrameStoreTests.MotionFieldsFollowAliasesAndPresentationOwnership|FullyQualifiedName~Av1ReferenceFrameStoreTests.MotionFieldAllocationFailureUnwindsTileReaderOwnership|FullyQualifiedName~Av1ReferenceFrameStoreTests.PartialReplacementPreservesSharedOwner|FullyQualifiedName~Av1ReferenceFrameStoreTests.FinalReplacementReleasesDisplacedOwner|FullyQualifiedName~Av1ReferenceFrameStoreTests.ResetReleasesUniqueOwnersAndClearsSlots|FullyQualifiedName~Av1ReferenceFrameStoreTests.TakeOutputTransfersPlanesAndReleasesOtherReferences|FullyQualifiedName~Av1ReconstructionConformanceTests.DecodeProgressiveSingleMatchesReference|FullyQualifiedName~Av1ReconstructionConformanceTests.DecodeProgressiveSingleWithConstrainedAllocator|FullyQualifiedName~Av1ReconstructionConformanceTests.DecodeProgressiveSingleTracksMotionFieldOwnership'
dotnet build src\ImageSharp\ImageSharp.csproj -c Release -f net10.0 --no-restore --disable-build-servers -m:1 --no-incremental --nologo --verbosity:minimal
dotnet build src\ImageSharp\ImageSharp.csproj -c Release -f net11.0 --no-restore --disable-build-servers -m:1 --no-incremental --nologo --verbosity:minimal
dotnet test tests\ImageSharp.Tests\ImageSharp.Tests.csproj -c Release -f net10.0 --no-restore --disable-build-servers -m:1 --filter $heifCheckpointFilter --logger 'console;verbosity=minimal'
dotnet test tests\ImageSharp.Tests\ImageSharp.Tests.csproj -c Release -f net11.0 --no-restore --disable-build-servers -m:1 --filter $heifCheckpointFilter --logger 'console;verbosity=minimal'
$aomVcVars = 'C:\Program Files (x86)\Microsoft Visual Studio\18\BuildTools\VC\Auxiliary\Build\vcvars64.bat'
$aomCmake = 'C:\Program Files (x86)\Microsoft Visual Studio\18\BuildTools\Common7\IDE\CommonExtensions\Microsoft\CMake\CMake\bin\cmake.exe'
$aomEnvironment = & cmd.exe /d /s /c "`"$aomVcVars`" >nul && set"
foreach ($aomEntry in $aomEnvironment)
{
$aomParts = $aomEntry -split '=', 2
if ($aomParts.Length -eq 2)
{
[Environment]::SetEnvironmentVariable($aomParts[0], $aomParts[1], 'Process')
}
}
& $aomCmake --build artifacts\reference\aom-generic --target aomdec --config Release --parallel 1
& 'artifacts\reference\aom-generic\aomdec.exe' --codec=av1 --rawvideo --all-layers --threads=1 --row-mt=0 --output='artifacts\reference\aom-generic\progressive-current-main-all-layers.yuv' 'tests\Images\Input\Heif\Av1\Conformance\libavif-progressive-draw-points-8b.bit'
3. Reverify downstream inter prediction in recorded order
The single-reference syntax, buffer, reconstruction, and ownership foundation is verified by 54bb6cbe59bd113058854a3ee31448cf61f462ca. Reverify the existing downstream implementations in this exact order, treating each as locally implemented but unverified until its current-main evidence is recorded.
- Compound reference selection, paired reference-MV derivation, and equal averaging.
- Inter-intra prediction.
- Distance-weighted compound prediction.
- Wedge compound prediction.
- Difference-weighted compound prediction.
- OBMC.
- Scaled-reference prediction.
- Local warped prediction.
- Non-translational global prediction.
- Inter deblocking decisions and reference/mode deltas.
Verified equal-average compound checkpoint evidence on 2026-08-31:
- Refreshed the clean official libaom
maincheckout and audited the observed revision441c439b9916474cac15d2822af47a9ad70674a8. Reference selection and compound mode syntax matchread_comp_reference_typeandread_ref_framesinav1/decoder/decodemv.c; contexts matchav1/common/pred_common.c; paired reference-MV construction and eight-entry extension matchprocess_compound_ref_mv_candidateandsetup_ref_mv_listinav1/common/mvref_common.c. - Audited equal-average reconstruction against
av1/common/convolve.candav1/common/convolve.h. Corrected the unscaled 10/12-bit translational path so both references retain libaom's no-round compound intermediates until the sole final average and clipping step, including the larger first-round shift required for 12-bit horizontal intermediates. - Added descending Vector512, Vector256, Vector128, and scalar high-bit-depth traversal to the existing semantic compound-prediction operator families. No per-block, per-row, or per-scanline allocation or copy was added.
- Added FeatureTestRunner coverage for 10/12-bit copy, horizontal, vertical, and separable subpixel prediction at widths 9, 17, 33, and 65, with an independent no-round bilinear oracle, row-padding sentinels, and explicit scalar comparison.
- Added a complete
Av1BlockDecoder.DecodeBlock10/12-bit half-sample regression whose expected result comes from the scalar no-round pipeline. The selected vector differs by one sample from the obsolete round-each-reference behavior, so the test proves the production branch selection. - Refreshed the official libaom
mainremote immediately before verification and decoded the fixture's 5,465-byte AV1mdatpayload with currentaomdec, one thread and row threading disabled. All 19 frames decoded; the final 19,200 YUV444 samples have SHA-256E79D2F49C260B1AC9B1B9BBBB2D611126AFD3B241DA389EB9E7BD4EA0ED42080and match the retained native reference with zero differing samples. - The real 19-frame production sequence requires decoded equal-average compound blocks, compares the final native Y, U, and V planes exactly, compares final RGBA presentation through ImageSharp's established reference-output API, and repeats the complete decode with a 1,024-byte constrained tracked allocator and exactly-once return checks.
- The focused Release checkpoint set passes 31/31 on net10.0 and 31/31 on net11.0, with zero
failures or skips. Scoped analyzer and whitespace verification pass for every changed C# file,
Roslynk reports zero compiler errors and no diagnostics in the changed files, and
git diff --checkpasses..gitattributesis unchanged. - The completed checkpoint was committed as
4075a0844836e863a93cb2e2f3ca42d202c7df1bwith author and committerJames Jackson-South <james_south@hotmail.com>.
Verified inter-intra checkpoint evidence on 2026-08-31:
- Audited syntax against current libaom
av1/decoder/decodemv.candav1/common/blockd.h. ImageSharp applies the same sequence enable, skip-mode, block-size, and single-reference gates, reads the same four-mode CDF, and reads wedge syntax only within libaom's wedge-supportedBLOCK_8X8throughBLOCK_32X32range. - Audited reconstruction against
ii_weights1d,ii_size_scales,build_smooth_interintra_mask, andcombine_interintrain currentav1/common/reconinter.c. The ImageSharp weights, plane-size scaling, smooth-mask direction, complemented destination orientation, wedge sign, subsampling, and final 6-bit blend match. No production change was required. - The mask tests cover all four inter-intra modes, complemented orientation, row-padding
sentinels, and the 32-wide curve. FeatureTestRunner covers byte and high-bit-depth selectable
blending under SIMD and scalar dispatch, and complete
Av1BlockDecoder.DecodeBlocktests execute smooth inter-intra reconstruction at 8, 10, and 12 bits. - Extracted the fixture's 5,327-byte AV1
mdatpayload and decoded it with the refreshed current libaomaomdec, using one thread with row threading disabled. All 19 frames decoded. The final 19,200 YUV444 samples have SHA-256E8B776C2751DC30CA838931A4B74535FC6E681179568A1278747A38CFF2E5BFAand match the retained native reference with zero differing samples. - The real production sequence requires both smooth and wedge inter-intra blocks, compares the final native Y, Cb, and Cr planes exactly, and compares final RGBA presentation through ImageSharp's established reference-output API. Its constrained 1,024-byte tracked-allocator run proves motion-field allocation and exactly one return for every allocation.
- The focused Release checkpoint set passes 50/50 on net10.0 and 50/50 on net11.0, with zero
failures or skips. Scoped analyzer and whitespace verification pass for both changed C# files.
Roslynk reports zero compiler errors and no diagnostics in the changed files,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
18b1c881271a3494489ca6f410ab140544902e2dwith author and committerJames Jackson-South <james_south@hotmail.com>.
Verified distance-weighted compound checkpoint evidence on 2026-08-31:
- Audited reference-distance quantization against
quant_dist_weightandquant_dist_lookup_tablein current libaomav1/common/common_data.h, and audited order-hint distance selection and forward/backward reference assignment againstav1_dist_wtd_comp_weight_assigninav1/common/reconinter.c. - Audited reconstruction against current libaom
av1/common/convolve.c. Corrected the production 10/12-bit subpixel path, which incorrectly finalized its two no-round compound intermediates with an equal average instead of the signaled distance weights. The fixed path applies libaom's 4-bit weighted shift before bias removal, final rounding, and clipping. - Added descending Vector512, Vector256, Vector128, and scalar traversal to the existing semantic distance-weighted intermediate predictor family. Unsigned widening preserves the biased 12-bit intermediate range. No per-block, per-row, or per-scanline allocation or copy was added.
- Added FeatureTestRunner coverage for every current-libaom distance-weight class in both reference orders, and for 10/12-bit copy, horizontal, vertical, and separable subpixel prediction at widths 9, 17, 33, and 65, with an independent no-round oracle and row-padding sentinels.
- Added a complete
Av1BlockDecoder.DecodeBlock10/12-bit half-sample regression that selects the 13:3 distance weights through real order hints. Its first reconstructed sample differs from the old equal-average result, so the test proves the corrected production branch is executed. - Extracted the fixture's 5,372-byte AV1
mdatpayload and decoded it with the refreshed current libaomaomdec, using one thread with row threading disabled. All 19 frames decoded. The final 19,200 YUV444 samples have SHA-256E8CAA650F1571C5B9CACAF8C06E1DDF5F5D2ED35F65F1C34377076C573425899and match the retained native reference with zero differing samples. - The real 19-frame production sequence requires decoded distance-weighted compound blocks, compares the final native Y, Cb, and Cr planes exactly, compares final RGBA presentation through ImageSharp's established reference-output API, and repeats the complete decode with a 1,024-byte constrained tracked allocator and exactly-once return checks.
- The focused Release checkpoint set passes 44/44 on net10.0 and 44/44 on net11.0, with zero failures
or skips. Scoped analyzer and whitespace verification pass for every changed C# file. Roslynk reports
zero compiler errors and no diagnostics in the changed files,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
7e2de7a2c25852acc374b17936a1a644464f77f3with author and committerJames Jackson-South <james_south@hotmail.com>.
Verified wedge compound checkpoint evidence on 2026-08-31:
- Audited mask generation against current libaom
tools/gen_wedge_masks_data.pyandav1/common/reconinter.c, including the master prototypes, direction transforms, block-size codebooks, sign flips, offsets, and luma/chroma mask sampling. ImageSharp's generated masks match those definitions; only stale “pinned” documentation required correction. - Audited reconstruction against current libaom
aom_dsp/blend_a64_mask.c. The high-bit-depth d16 path applies the Q6 mask to both no-round intermediates before bias removal, the sole final rounding step, and clipping. - Corrected the production high-bit-depth intermediate eligibility gate, which admitted only equal-average blocks and made the distance-weighted and wedge no-round finalizers unreachable. Average, distance-weighted, and wedge subpixel blocks now retain both intermediates until their signaled finalizer; difference-weighted blending remains excluded for its next ordered checkpoint.
- Added high-bit-depth traversal to the existing semantic mask-blend predictor and readonly operator family with descending Vector512, Vector256, Vector128, and scalar dispatch. Unsigned widening preserves the biased 12-bit intermediate range. No per-block, per-row, or per-scanline allocation or copy was added.
- Extended FeatureTestRunner coverage with an independent Q6 mask oracle across 10/12-bit copy,
horizontal, vertical, and separable subpixel prediction, widths 9, 17, 33, and 65, all mask weights
from 0 through 64, and row-padding sentinels. A complete
Av1BlockDecoder.DecodeBlockregression verifies the current-libaom 8x8 wedge mask and the production no-round branch. - Extracted the fixture's 5,374-byte AV1
mdatpayload and decoded it with refreshed current libaomaomdec, using one thread with row threading disabled. All 19 frames decoded. The final 19,200 YUV444 samples have SHA-256E8CAA650F1571C5B9CACAF8C06E1DDF5F5D2ED35F65F1C34377076C573425899and match the retained native reference with zero differing samples. - The real 19-frame production sequence requires both wedge-mask orientations, compares final native Y, Cb, and Cr planes exactly, compares final RGBA presentation through ImageSharp's established reference-output API, and repeats the complete decode with a 1,024-byte constrained tracked allocator and exactly-once return checks.
- The focused Release checkpoint set passes 35/35 on net10.0 and 35/35 on net11.0, with zero
failures or skips. Scoped analyzer and whitespace verification pass for every changed C# file.
Roslynk reports zero compiler errors,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
9883a24dc319e16b471f68be632d4f62f2c1cd5ewith author and committerJames Jackson-South <james_south@hotmail.com>.
Verified difference-weighted compound checkpoint evidence on 2026-08-31:
- Audited syntax against current libaom
av1/decoder/decodemv.c. ImageSharp applies the same masked-compound enable and block-size gates, selects difference-weighted compound directly when wedge is unavailable, and reads the same one-bit type-38 mask orientation. - Audited mask generation and reconstruction against current libaom
av1/common/reconinter.candaom_dsp/blend_a64_mask.c. The d16 path rounds the absolute intermediate difference by the convolution and bit-depth shift, scales it by 1/16, adds the type-38 base, clamps or inverts the mask, and then blends the original no-round intermediates before final rounding and clipping. Chroma reuses the luma-derived mask through rounded subsampling. - Corrected the production 10/12-bit subpixel eligibility gate, which previously rounded both references before difference-mask construction and blending. Difference-weighted blocks now use the existing semantic intermediate mask-builder and mask-blend predictor/operator families through the sole final rounding step. No new operator family, per-block allocation, or copy was introduced.
- Renamed the stale “pinned formula” test and extended FeatureTestRunner's independent oracle across current-libaom regular and d16 mask arithmetic, both mask orientations, 8/10/12-bit samples, widths that cross every Vector512, Vector256, Vector128, and scalar boundary, subpixel phases, and row-padding sentinels.
- Added a complete
Av1BlockDecoder.DecodeBlockregression for 10/12-bit half-sample prediction and both type-38 orientations. Its expected mask and reconstruction are calculated directly from the current-libaom equations, independently of the production mask builder and finalizer. - Extracted the fixture's 5,358-byte AV1
mdatpayload and decoded it with refreshed current libaomaomdec, using one thread with row threading disabled. All 19 frames decoded. The final 19,200 YUV444 samples have SHA-256E8CAA650F1571C5B9CACAF8C06E1DDF5F5D2ED35F65F1C34377076C573425899and match the retained native reference with zero differing samples. - The real 19-frame production sequence requires both difference-mask orientations, compares final native Y, Cb, and Cr planes exactly, compares final RGBA presentation through ImageSharp's established reference-output API, and repeats the complete decode with a 1,024-byte constrained tracked allocator and exactly-once return checks.
- The focused Release checkpoint set passes 37/37 on net10.0 and 37/37 on net11.0, with zero
failures or skips. Scoped analyzer and whitespace verification pass for every changed C# file.
Roslynk reports zero compiler errors,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
fb4c64474e1ced4067a42731384f3b5ad4212a2fwith author and committerJames Jackson-South <james_south@hotmail.com>.
Verified OBMC checkpoint evidence on 2026-08-31:
- Audited motion-mode syntax against current libaom
av1/decoder/decodemv.c,av1/common/blockd.h,av1/common/reconinter.c,av1/common/obmc.h, andav1/common/reconinter_template.inc. ImageSharp applies the same switchable-mode, skip, single-reference, inter-intra, minimum-size, overlappable-neighbor, fixed-global-motion, scaled reference, and projection-sample gates and reads the matching binary or three-way CDF. - Audited above and left neighbor traversal, 4x4 pairing, neighbor caps, chroma suppression, prediction rectangles, interpolation filters, first-reference selection, mask tables, and blend order against current libaom. The existing semantic mask-blend predictor remains the correct SIMD-first traversal; no OBMC-specific operator family, allocation, or copy was introduced.
- Corrected the unscaled neighbor far-edge UMV clamp. After converting libaom's neighbor-relative motion-vector limits to an absolute source coordinate, the prediction extent cancels from the right and bottom limits; the previous code counted it twice.
- Extracted the fixture's 5,387-byte AV1
mdatpayload at AVIF offset 1,065 and decoded it with refreshed current libaomaomdec, using one thread with row threading disabled. All 19 frames decoded. The final 19,200 YUV444 samples have SHA-256E8CAA650F1571C5B9CACAF8C06E1DDF5F5D2ED35F65F1C34377076C573425899and match the retained native reference with zero differing samples. - The production sequence asserts decoded OBMC mode state, compares final native Y, Cb, and Cr
planes exactly, compares final RGBA presentation through ImageSharp's established reference-output
API under normal and scalar FeatureTestRunner dispatch, and repeats reconstruction with a 1,024-byte
constrained tracked allocator. Direct
DecodeBlocktests cover above-then-left blending at 8/10/12-bit and 4:2:0 and 4:2:2 chroma geometry. - Renamed the stale pinned-reference test and its established reference-output PNG together. The
PNG SHA-256 remains
D2CB388C9092EF17C4F0382C0150DD30D6F9D0EE247FF45AB5D7D4D312CEB23C; only its contract-derived filename changed. - The focused Release checkpoint set passes 18/18 on net10.0 and 18/18 on net11.0, with zero
failures or skips. Scoped analyzer and whitespace verification pass for every changed C# file.
Roslynk reports zero compiler errors,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
7e7e3cbe6438d63926b31d966795d2652e221939with author and committerJames Jackson-South <james_south@hotmail.com>.
Verified scaled-reference checkpoint evidence on 2026-08-31:
- Audited reference-size validation and variable-scale coordinates, filters, edge extension, convolution
rounding, and compound intermediates against current libaom
av1/common/scale.c,av1/decoder/decodeframe.c, andav1/common/convolve.c. The frame boundary accepts the same half-to-sixteen-times dimension range and requires at least one compatible selected reference. - Corrected the production scaled-compound branch. It previously rounded each scaled reference into
native pixels before blending; current libaom retains both
CONV_BUF_TYPEvalues withCOMPOUND_ROUND1_BITSequal to seven and performs one final rounding after the selected compound blend. - Kept native-pixel and compound output in the existing
Av1ScaledInterPredictortraversal with semanticNativeOperatorandCompoundOperatoroutput contracts. The closed generic traversal shares variable-phase arithmetic across byte and ushort sources, dispatches Vector512, Vector256, Vector128, then scalar, and adds no per-block allocation or copy. - Added independent FeatureTestRunner oracles for native and no-round compound output across 8, 10,
and 12 bits, variable phases, all interpolation families, reduced kernels, vector tails, and destination
padding. A complete
Av1BlockDecoder.DecodeBlock()regression covers scaled compound prediction across all, AVX-512-disabled, AVX-disabled, and scalar configurations and proves the vector differs from an incorrectly early-rounded blend. - Decoded the 2,195-byte layered payload with refreshed current libaom
aomdec, using one thread, row threading disabled, all layers selected, and raw 8-bit output. The 40x40 YUV444 base and 80x80 YUV444 dependent frames total 24,000 samples with SHA-256DD219E41B52C6C9343A92CD0A2D451DF57B73B25F10124811675B4CB2F8D666F; both match their retained native references with zero differing samples. - The production tests compare both native frames exactly, compare selected-layer and final RGBA presentation through ImageSharp's established reference-output API, and repeat both paths with a 1,024-byte constrained tracked allocator whose allocations have balanced exactly-once returns.
- Renamed the two stale pinned-reference tests and their contract-derived PNGs together. Their Git blob
identifiers remain unchanged, and their SHA-256 values remain
DC4C6DBE6BD92C5FCE1E3E23700AFA603EF04ED02EDD336213EBBA1E3BD84BA0and678C5E5D4650EA6F0C590302E7DB9E3C6608851BC577453DA4A6837BDB4D3AF3. - The focused Release checkpoint set passes 10/10 on net10.0 and 10/10 on net11.0, with zero failures
or skips. Scoped analyzer and whitespace verification pass for every changed C# file. Roslynk reports
zero compiler errors,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
658a9cd1b6e22806decbae923da8800bca03a09ewith author and committerJames Jackson-South <james_south@hotmail.com>.
Verified local warped-prediction checkpoint evidence on 2026-08-31:
- Refreshed the clean official libaom
maincheckout and audited the observed revision441c439b9916474cac15d2822af47a9ad70674a8. Motion-mode eligibility and CDF selection matchread_motion_modeinav1/decoder/decodemv.c; above, left, top-left, and top-right spatial projection samples and threshold selection matchfindSamplesandselectSamplesinav1/common/mvref_common.c; affine fitting, shear reduction, phase derivation, filters, rounding, clipping, and invalid-model fallback matchav1/common/warped_motion.candav1/common/reconinter.c. - Mechanically compared all 1,544 ImageSharp and independent-test warped-filter coefficients against
current libaom's
av1_warped_filter; both comparisons have zero differences. The separate scalar test transcription covers 8-, 10-, and 12-bit luma and subsampled-chroma coordinates, tail widths, destination stride preservation, libaom's 12-bit round adjustment, and AVX-512, AVX, 128-bit, and scalar dispatch throughFeatureTestRunner. - Extracted the fixture's exact 2,310-byte AV1
mdatpayload at AVIF offset 997. Its SHA-256 is644D04FE1D1A32BB7A3856AD7EB49CF1EFDE0AC845E55BEAC4170F72353F2391. Current official libaom decoded both 256x256 YUV444 frames with one thread, row threading disabled, and all layers enabled. The complete Y4M SHA-256 is8FDC5D46014F5E5A7455A83643AB6F0DA66FC5A984E72A43F8C75BAD8271C299; the final frame's 196,608 native samples have SHA-25647B2AB39BF3B9DA15C1EC59840F964DFDF227760947F6E1295FB38A84555F75Cand match the retained native reference with zero differences. - The real two-frame fixture exercises
Av1BlockDecoder.DecodeBlock(), requires decodedWARPED_CAUSALstate and the expected multi-sample affine model, compares final native Y, U, and V planes exactly, compares the retained final presentation through ImageSharp's established reference-output API, and passes through intrinsic and scalar dispatch. The 1,024-byte constrained tracked-allocator path passes with motion-field allocations present and balanced exactly-once returns. - Renamed the stale pinned-reference test and its contract-derived PNG together without changing the PNG
bytes. Its SHA-256 remains
4490D62FB6679378E92CACA48427359091AD2106BE49FC1A3848F78BE03BEEB1. - The focused Release checkpoint set passes 4/4 on net10.0 and 4/4 on net11.0, with zero failures or
skips. Scoped analyzer verification passes for both changed C# files. Roslynk reports zero compiler errors,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
27a522424fe7aaea25078e705d71a501da110727with author and committerJames Jackson-South <james_south@hotmail.com>.
Verified non-translational global-prediction checkpoint evidence on 2026-08-31:
- Audited global-motion syntax, coefficient decoding, previous-reference recentering, shear validation,
motion-vector projection, and warped-prediction eligibility against current official libaom
mainat the observed revision441c439b9916474cac15d2822af47a9ad70674a8. The implementation matchesread_global_motion_params,read_global_motion_model,gm_get_motion_vector,is_global_mv_block, and the WARP_PRED selection inav1/common/reconinter.c. - Corrected high-bit-depth compound warped/global prediction to retain both references in libaom's
unsigned no-round compound domain. Current
get_conv_params_no_round,av1_warp_plane, andav1_highbd_warp_affine_crequire the 12-bit first-round adjustment while retaining a seven-bit second round; native clipping now occurs only after the compound blend. - The independent scalar libaom transcription validates native and no-round compound output for byte,
8-bit, 10-bit, and 12-bit sources, including tail widths and destination-stride preservation. All cases pass
through AVX-512, AVX, 128-bit, and scalar dispatch with
FeatureTestRunner. DirectAv1BlockDecoder.DecodeBlock()coverage validatesGLOBAL_GLOBALMVcompound reconstruction at all supported bit depths. - Extracted the fixture's exact 38,475-byte AV1
mdatpayload at AVIF offset 997. Its SHA-256 is6AC7EC9984B1FF5C00403D7E3858441E9CEE75128F7414101D06DEEE59A351D0. Current official libaom decoded both 256x256 YUV444 frames with one thread, row threading disabled, and all layers enabled. The complete Y4M SHA-256 is84754DE0B9FABC4F3F8F344C848183EC17B625BFD87E4519C3D8AD7DEFD20F2C; the final frame's 196,608 native samples have SHA-256FEC89E2DE7496980389806B194425042F3800C7BAA817249D1A51D44A2B37A8Eand match the retained native reference with zero differences. - The real two-frame fixture exercises the production decoder, requires decoded non-translational global motion, compares final native Y, U, and V planes exactly, compares the retained presentation through ImageSharp's established reference-output API, and passes the constrained tracked-allocator path.
- Renamed the stale pinned-reference test and its contract-derived PNG together without changing the PNG
bytes. Its SHA-256 remains
F7D27ABF79450DFA311F72106FD1DA80997EABC0937F2F5578EF627119FF83B0, and Git attributes select the LFS filter and diff driver. - The focused Release checkpoint set passes 11/11 on net10.0 and 11/11 on net11.0, with zero failures or
skips. Scoped analyzer verification passes for all six changed C# files. Roslynk reports zero compiler
errors,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
25295683d39a2336e9b98484c9fd54f33107ea66with author and committerJames Jackson-South <james_south@hotmail.com>.
Verified inter-deblocking checkpoint evidence on 2026-08-31:
- Audited frame-level loop-filter syntax and primary-reference inheritance against
setup_loopfilterin currentav1/decoder/decodeframe.c; per-superblock delta-LF parsing and prediction againstread_delta_q_paramsinav1/decoder/decodemv.c; and default reference/mode deltas againstav1/common/entropymode.cat observed current-main revision441c439b9916474cac15d2822af47a9ad70674a8. - Audited filter-level derivation, segmentation adjustment, reference scaling, global/non-global
mode classes, skipped-transform prediction-unit decisions, transform-edge selection, kernel length,
sharpness limits, and vertical-then-horizontal traversal against
get_filter_level,set_lpf_parameters,av1_filter_block_plane_vert,av1_filter_block_plane_horz, andav1_thread_loop_filter_rows. No production arithmetic change was required. - Added direct production
Av1LoopFilterDecoder.DecodeFrame()coverage using adjacent skipped 16x8 inter blocks split into 8x8 transforms. An independent scalar oracle proves that internal transform edges remain untouched and the prediction-unit edge uses current-libaom levels 17 for LAST/GLOBALMV, 21 for LAST/NEWMV, and 22 for GOLDEN/GLOBALMV. ExistingFeatureTestRunnercoverage continues to verify every filter width at 8, 10, and 12 bits under intrinsic and scalar dispatch. - Current official libaom decoded the retained 20,750-byte 8-bit, 37,169-byte 10-bit, and
23,769-byte 12-bit elementary streams with one thread, row threading disabled, raw output, and their
native output depths. The generated native files match the retained references byte for byte. Their
output SHA-256 values are
8DDE2EEC742C39F0579C29AE84CBA0FE01522A9008ADCB2CFFCCEC0295D18141,9A59DD92A0C579F942ACCA8281EBD0465DC848BE200A4D2FF57EAFF589445F6C, andEF712BE32AF7CF0A95C5C41BDCC51AFC05A4AB7C047383F5F65EDAD2BB986712. - Reused the already current-main scaled-reference sequence as the real inter checkpoint. It requires an inter frame with reference/mode-delta processing enabled, nonzero chroma filter levels, intra, inter, and skipped-inter blocks; compares both decoded native frames exactly; compares final presentation through ImageSharp's established reference-output API; and passes constrained tracked allocation with balanced returns.
- Removed an obsolete SVT-AV1 design link from mode-map documentation. Current official libaom remains the sole external codec implementation source.
- The focused Release checkpoint set passes 6/6 on net10.0 and 6/6 on net11.0, with zero failures or skips.
- Scoped analyzer verification passes for all four changed C# files. Roslynk reports zero compiler
errors,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
fcb502e4960cc7b8efb06b6f060e2c73a913a2bfwith author and committerJames Jackson-South <james_south@hotmail.com>.
For every item:
- Trace syntax and arithmetic to the current libaom
maintree. - Execute the real production decoder path.
- Compare native planes exactly.
- Compare presentation through the established reference-image API.
- Run constrained allocator and exactly-once ownership coverage.
- Run FeatureTestRunner for SIMD and scalar dispatch when the implementation has SIMD.
- Record focused Release evidence before marking the item verified.
4. Close AV1 decoder coverage
Previously verified algorithm checkpoints remain valuable evidence, but the final decoder gate requires a fresh current-tree run after the inter and cleanup corrections.
- Bounded OBU framing, sequence headers, frame headers, tile groups, alignment, and trailing-bit parsing have been re-audited and verified against current libaom
main. - Partition traversal, mode information, segmentation, delta quantization, transform-size selection, coefficient decoding, inverse quantization, and inverse transforms have been re-audited and verified against current libaom
main. - Intra prediction covers directional, DC, smooth, Paeth, chroma-from-luma, filter-intra, and palette families with the established operator architecture.
- Intra-block copy has exact native reconstruction and feature-isolated SIMD evidence.
- Lossless inverse transform, loop filtering, CDEF, super-resolution, restoration, and film grain have focused checkpoint evidence.
- Retained references, CDF snapshots, segmentation maps, global motion, temporal motion fields, and dependent-frame lifecycle have been re-audited and verified against current libaom
main. - The 12-case all-intra profile matrix covers every valid 8, 10, and 12-bit monochrome, 4:2:0, 4:2:2, and 4:4:4 combination. Dependent-frame coverage is recorded separately above.
- The exact current-tree native-plane matrix passes through the production decoder on net10.0 and net11.0. The normal-dispatch and FeatureTestRunner fallback methods pass 2 of 2 focused tests on each target.
- The exact current-tree presentation matrix passes 12 of 12 cases through ImageSharp's established reference-image API on net10.0 and net11.0.
- Verify malformed/truncated data, frame IDs, reference slots, tile bounds, allocation limits, cancellation, and failure unwinding.
- Verify still items and bounded sequences from file, memory, non-seekable, and short-read streams.
- [~] Verify ICC, CICP, alpha, grids, pixel aspect ratio, clean aperture, rotation, mirroring, metadata, and every presented sequence frame. Grid validation now requires the first cell to be at least 64 samples on both axes, enforces even output and cell dimensions along each subsampled AV1 chroma axis, and requires every cell to cover its row-major output region without exceeding the first cell's dimensions. This accepts the smaller right and bottom cells supported by the writer while also accepting uniform coded cells whose final row and column are cropped to the grid descriptor. Auxiliary alpha uses the same cropped overlap, so padded cells cannot write outside the final frame. Production regressions cover smaller right, bottom, and bottom-right color and alpha cells; Roslynk compiler and scoped analyzer diagnostics are clean, while runtime verification of the current tree remains pending.
- Complete the public AVIF format/API review so registered capabilities match implemented behavior.
- Remove or reject every valid in-scope AV1 syntax branch that remains silently ignored or unsupported.
Verified negative-path and frame-identifier gate evidence on 2026-08-31:
- A two-frame lossless frame-identifier sequence was generated and decoded with the clean official
libaom
maincheckout at observed revision441c439b9916474cac15d2822af47a9ad70674a8. Both decoded frames match the source Y, Cb, and Cr samples exactly. DecodeFrameIdentifiersMatchReferenceexecutes the production decoder through FeatureTestRunner, compares both native frames exactly, and proves the second frame is dependent with a changed current frame identifier. The current-frame, reference-delta, stale-slot, and refreshed-slot identifier logic was audited against the same currentmainsource.- The focused negative-path set passes 46 of 46 cases on net10.0 and 46 of 46 on net11.0, with zero failures or skips. It covers truncated palette entropy, malformed-following-OBU recovery, parser lifecycle failure, overflowing and invalid tile bounds, reference-slot ownership and transfer, constrained multi-group allocation, motion-field allocation failure unwinding, and frame identifiers.
- The established paused-stream cancellation suite now includes AVIF. It verifies cancellation at 0%, 30%, and 70% of both file and memory streams, plus pre-cancelled identification, on both targets.
- The completed checkpoint was committed as
7f0e08126b3354e8f1eb45886f0d572006ae27dewith author and committerJames Jackson-South <james_south@hotmail.com>.
Verified bounded-OBU checkpoint evidence on 2026-08-31:
- Audited
av1/decoder/obu.c,av1/decoder/decodeframe.c,av1/common/obu_util.c,av1/common/tile_common.c,aom/src/aom_integer.c, andaom_dsp/bitreader_buffer.cin the clean official libaommaincheckout. BothHEADandorigin/mainresolved to the observed revision441c439b9916474cac15d2822af47a9ad70674a8; this is verification evidence, not a pin. - The bounded container scanner and production OBU reader now agree with current libaom on ignored reserved header fields and the shared unsigned 32-bit LEB128 limit.
- Sequence-header validation now rejects undefined level indices, initial display delays above ten, frame identifiers above sixteen bits, zero timing units, the UVLC overflow sentinel, and invalid identity-matrix profile or subsampling combinations at the owning syntax boundary.
- Frame and tile parsing now rejects
show_existing_framein a combinedOBU_FRAME, the all-slots intra-only refresh mask, inner tile columns below current libaom's super-resolution-aware minimum, overflowing or out-of-bounds tile sizes, and empty final tile payloads. - The still-image writer now emits the required zero tile-bound-presence bit for a multi-tile combined
OBU_FRAME, matching current libaom's single-tile-group encoder path. ObuFrameHeaderTestsandObuFrameLifecycleTestscover the corrected syntax through the real bounded parser. The focused parser set passes 50 of 50 cases on net10.0.- The final focused production set passes 55 of 55 cases on net10.0 and 55 of 55 on net11.0, with zero failures or skips. It includes exact final-layer and selected-layer native planes, exact established reference-image presentation, constrained allocator ownership, malformed-following-OBU recovery, and FeatureTestRunner normal, AVX-512-disabled, AVX-disabled, and scalar execution.
- A fresh direct foreground current-main
aomdecrun decoded both progressive layers with one thread and row multithreading disabled. All 2,178 Y, U, and V samples match the retained YUV444-alpha reference; the alpha plane is excluded from the AV1 native-plane comparison. - The current-libaom production reference test and its established PNG were renamed together. The PNG
bytes remain unchanged at SHA-256
0758C17DC36E38AEE9F4389A335C2BF332AB91E4C79D7B0B22994FDDD0FD1605, both paths resolve todiff=lfs, and.gitattributeswas not edited. - Release source builds pass for net10.0 and net11.0 with zero warnings and zero errors. Roslynk reports zero compiler errors, and scoped production and test analyzer verification reports no changes.
- The completed checkpoint was committed as
243524c2c0b52a49d8d161fab806ab092cabe47cwith author and committerJames Jackson-South <james_south@hotmail.com>.
Verified partition, mode, segmentation, quantization, and transform checkpoint evidence on 2026-08-31:
- Audited partition traversal and chroma representability against
read_partitionand the subsampled plane-size rejection in current libaomav1/decoder/decodeframe.c; spatial segment-ID decoding and corruption handling againstread_segment_idinav1/decoder/decodemv.c; delta-Q syntax, resolution, arithmetic, and clamping againstread_delta_qindexandread_delta_q_paramsin the same file. - Audited selected and variable transform-size traversal against
read_tx_size,read_tx_size_vartx, and transform-block traversal inav1/decoder/decodeframe.c; coefficient syntax and arithmetic againstav1_read_coeffs_txbinav1/decoder/decodetxb.c; inverse quantization and transform application against currentav1/decoder/decodeframe.c,av1/common/idct.c, and the current libaom transform test oracle. The observed cleanHEADandorigin/mainrevision was441c439b9916474cac15d2822af47a9ad70674a8; this is verification evidence, not a pin. - Partition decoding now rejects an invalid partition subsize and a block size that cannot represent the current subsampled chroma plane. Spatial segmentation rejects decoded IDs above the active segment range. Focused tests exercise both current-libaom corruption boundaries through the production tile reader.
- Coefficient entropy decoding uses one allocator-owned maximum-size
Av1LevelBufferper tile reader. Each transform resets and clears only its active padded geometry, so no transform creates an allocation. Allocation tracking over all eight minimum- and maximum-quantizer frames proves exactly one coefficient scratch allocation per frame and exactly-once return after decoder disposal. - Palette index maps use one allocator-backed 32 KiB decoder-session owner with non-owning 128x128 luma
and chroma views. Each parsed superblock is reconstructed before either view is reused, and each block
clears only its transient
Buffer2DRegionafter prediction. The fixed session cost replaces the former full-frame maps without copies, fragmented memory groups, constructor rollback, or per-block allocations. The one-rent ownership regression and native palette reconstruction pass on net11.0, the complete HEIF/AV1 namespace passes 8,808 of 8,808 direct VSTest cases in Release, and the four-case palette set passes with exact native and presentation output, truncated-entropy rejection, and balanced exactly-once disposal. Av1BlockModeInfois value storage, removing the managed object allocation formerly created for every decoded coding block. ExplicitModeInfoIndexvalues preserve libaom's mode-info identity semantics at prediction-unit loop-filter edges, and the frame map now uses integer offsets so more than 65,535 decoded blocks cannot wrap its lookup identity.- Current official libaom reproduced the 39-frame all-intra reference and all four 8/10-bit minimum- and
maximum-quantizer references byte for byte. The production tests compare every native sample exactly,
cover every intra mode and seven selected transform types, execute SIMD and scalar paths through
FeatureTestRunner, and exercise the quantizer sequences under constrained tracked allocation. - Current official libaom decoded the 42-byte palette payload into the retained 1,089-byte YUV444
reference at SHA-256
E05F7C0DF06ECCF0E43869D1D7B03DAA1D635ACD26A766F8940899BE18D53251. The exact native test requires luma and chroma palette syntax. The established reference-output test uses the unchanged presentation PNG at SHA-2561148EBF6AA4B0F2D069D5E9B9605F6FB2A315E525F18016CDCAE23EFDD81DA84, whose renamed path still resolves todiff=lfs;.gitattributeswas not edited. - The exact final AV1 namespace passes 8,732 of 8,732 cases on net10.0 and 8,732 of 8,732 cases on net11.0, with zero failures or skips. Release source builds pass for net10.0 and net11.0 with zero warnings and zero errors. Roslynk reports zero compiler errors, and scoped analyzer verification reports no changes.
- The completed checkpoint was committed as
57a3f6668e39d0934e7b6b8d37a3dc2a5adc88f0with author and committerJames Jackson-South <james_south@hotmail.com>.
Verified retained-frame lifecycle checkpoint evidence on 2026-08-31:
- Audited primary-reference entropy selection, independent per-tile CDF starts, context-update-tile
publication, segmentation-map inheritance, reference-map refresh, and show-existing key-frame reset
against current libaom
av1/decoder/decodeframe.c,av1/decoder/decodemv.c,av1/decoder/decoder.c, andav1/common/entropymode.c. - Audited retained motion-vector cells, reference-side classification, projection source ordering,
projection limits, and reference-frame publication against
av1_copy_frame_mvs,av1_calculate_ref_frame_side,motion_field_projection, andav1_setup_motion_fieldin current libaom. Same-role primary-reference global-motion inheritance remains covered by the exact current-main global-warp fixture. The observed cleanHEADandorigin/mainrevision was441c439b9916474cac15d2822af47a9ad70674a8; this is verification evidence, not a pin. - Current official libaom decoded the retained
cdfupdate,mfmv,svc-L2T1,svc-L1T2, andsvc-L2T2streams with one thread, row threading disabled, and eight-bit output depth. Their generated Y4M files match the retained references byte for byte at SHA-2564FBFF73FF0DE2D9084DAE557D1D4BD677B0486516525BF4D327D2D795D5A7779,F7DB607694818C19E62FD9A27F53E1A3E2D00B72C39C0430C1B26399CC76777D,7A427631ECBF144F435AA4612F1201415FB1A9BCF9A67BA010AEF830B0C3AB81,4012DE2D4AFD095E7BB68EAE18B50B0674781BB4971CECABC0E5471E63373ED3, and1ABB981CFF76BA9557DA437B258D8A95FCA755DED8E3949D857E8388AB1D6AE3. - Existing allocation-tracking tests exercise initialization, retained-slot aliases, allocation-failure unwinding, presentation ownership, decoder-result ownership, repeated disposal, and final exactly-once return of reference frames, frame-owned motion fields, entropy snapshots, and segmentation maps.
- The focused Release checkpoint set passes 54 of 54 cases on net10.0 and 54 of 54 cases on net11.0, with zero failures or skips. It includes exact native CDF-update, motion-field, spatial-layer, temporal-layer, spatial-temporal-layer, progressive dependent-frame, and global-warp production paths, plus constrained allocator coverage.
- Release source builds pass for net10.0 and net11.0 with zero warnings and zero errors. Roslynk
reports zero compiler errors, scoped analyzer verification reports no changes,
git diff --checkpasses, and.gitattributesis unchanged.
Final decoder allocation, lifetime, precision, architecture, and test-validity audit evidence on 2026-09-01:
- Refreshed the official libaom remote and audited against observed
origin/main976867526367f571a1c09b994066af8364aed781. The intervening external-rate-controller commit does not changeav1/decoder,av1/common,aom_dsp, or the AV1 decoder build definition. - CDEF now uses one bounded 64x64-unit bordered source workspace, two preserved top-row slots per plane, preserved left columns, and unit-local direction and variance storage. This replaces the frame-wide source copy and frame-wide direction maps while retaining libaom's unit traversal and cross-plane luma-direction lifetime.
- Loop restoration now retains the required immutable source and separate destination, but stores the
full destination in native sample width. Eight-bit filtering narrows only bounded unit output after
clipping, while high-bit-depth filtering writes directly to the native
ushortdestination. - Reference-to-presentation copying now copies visible native rows only. Padding remains destination owned, and the ownership tests mutate a copied visible sample rather than unrelated padding.
- The remaining decoder allocations and copies are either bounded scratch or required ownership boundaries. Frame planes enforce their contiguous single-span invariant before allocation; palette, transform, film-grain, super-resolution, color-conversion, and alpha workspaces remain bounded and allocator owned. No per-block managed allocation remains in reconstruction.
- Block reconstruction now uses one exact-size signed-short owner for inverse quantization, inverse transform, compound prediction, convolution, and chroma-from-luma scratch. Even-length slices provide the integer workspaces without another rent. Monochrome reserves no chroma coefficients, and 4:2:0, 4:2:2, and 4:4:4 reserve two symmetric chroma planes at their coded subsampling. This replaces three constructor rents and their catch-all rollback path; exact allocation length, coefficient span length, and exactly-once return pass for all four layouts, with 549 adjacent reconstruction tests passing direct net11 VSTest in Release.
- Valid unsupported tile-list syntax is rejected explicitly. Reserved and metadata OBUs are consumed only after bounded framing and trailing-bit validation. Eight-, ten-, and twelve-bit reconstruction, presentation, alpha, restoration, and film-grain paths retain native precision.
- Predictor traversal remains split into semantic readonly operator families. The planar sample
adapter and transform-block context are value types, and Release construction sites use
defaultwithout null-forgiving suppression. - The net11.0 Release test project builds with zero errors. Roslynk reports zero compiler errors,
git diff --checkpasses, and.gitattributesis unchanged. - Visual Studio 18.9 VSTest ran the complete
Formats.Heif.Av1namespace with collection parallelism disabled and stop-on-failure enabled: 8,746 of 8,746 cases passed. The touchedHeifDecoderTestsandHeifSequenceParserTestsadd 104 of 104 passing integration cases. Focused CDEF, restoration, film-grain, copy-ownership, and reference-isolation runs also pass 15 of 15 cases. The historical report recorded successful process exits; it did not independently establish absence of Windows application-error dialogs.
Final decoder stream, presentation, and public-registration evidence on 2026-09-01:
- Real AV1 still-item and timed-sequence files decode identically from a file stream, memory stream, non-seekable stream, and a seekable stream limited to three bytes per read. All eight stream rows pass through public format detection and production decoding, comparing every presented frame exactly.
- A two-frame production sequence applies a centered clean-aperture crop, counter-clockwise rotation, mirroring, pixel-aspect-ratio metadata, and CICP metadata to every frame. The complete five-frame real auxiliary-alpha sequence composes non-opaque alpha and retains timing, Exif, and XMP for every frame.
- The fixed-header detector accepts both compact and extended-size leading file-type boxes. Default configuration registers the implemented HEIF decoder and detector but no longer advertises the incomplete HEIF encoder.
- Visual Studio 18.9 VSTest, serialized with stop-on-failure enabled, passes the 12 of 12 new
stream/presentation/registration cases and the complete current
HeifDecoderTestsplusHeifSequenceParserTestsset with the registration contract: 115 of 115. The final explicit no-encoder registration assertion passes 1 of 1 after its final edit. - The net11.0 Release test project builds with zero errors, Roslynk reports zero compiler errors,
git diff --checkpasses, and.gitattributesis unchanged. The historical report recorded normal VSTest exits and no surviving test host; that does not establish absence of Windows application-error dialogs.
SIMD traversal consistency evidence on 2026-09-02:
- The shared
Numericsvector-count helpers now cover same-lane spans and all fixed hardware widths. AV1 decoder and current encoder hot paths use those helpers for complete-vector traversal instead of repeating local modulo or last-vector calculations. Reverse-source indexing and algorithm-specific partial-output groups remain explicit because they are not vector-count calculations. - Forward quantization, palette prediction, and scaled inter prediction construct width-specific SIMD constants only when at least one vector batch will execute. Narrower dispatch tiers consume only the remainder left by wider tiers before the scalar tail.
- The net11.0 Release production assembly builds with zero warnings and zero errors. The test project
builds with zero errors while retaining the existing repository warning set. Roslynk reports zero
compiler errors,
git diff --checkpasses, and.gitattributesis unchanged. Foreground VSTest passes 63 of 63 focused quantizer, forward-transform, CDEF, restoration, palette, intra, inter, film-grain, and super-resolution cases.
Decoder exit gate:
- Re-establish current-main native evidence for every supported AV1 tool through the complete production path.
- Re-establish presentation behavior against independent reference images at the correct output precision.
- Complete the decoder audit for managed execution, copies, per-block allocation, and segmented memory.
- Verify allocator ownership and exceptional-path disposal throughout the decoder.
- Record focused Release verification of the final tree; earlier checkpoint results do not close these gates.
AV1 encoder implementation
Writer primitives are not an encoder. The public encoder remains incomplete until its complete decision and reconstruction paths follow the reference, every exposed option has production-path verification, and separately encoded output satisfies the one-unit per-sample acceptance limit.
5. Define and enforce the encoder contract
- Use official libaom
mainatd565eec60f084421fa34fc0534b760c6452b6a6cas the encoder syntax, probability-model, transform, quantization, filtering, and bitstream reference. - Use the existing PNG, TIFF, and JPEG encoders as the ImageSharp architecture reference: generic
Image<TPixel>input, encoder options taking precedence over converted format metadata and codec defaults, allocator-owned temporary storage, and deterministic disposal. - Treat source pixel type, source alpha representation, and decoded source bit depth as conversion inputs, never as output-eligibility checks. Do not pre-scan pixels before encoding.
- Resolve output configuration once from explicit encoder options, converted
HeifMetadata, and AV1 defaults in that order. Sanitize only combinations that cannot describe a legal requested output, and never write resolved values back to source metadata. - [~] Finalize observable options for quality, effort, lossless mode, bit depth, chroma subsampling, alpha quality, metadata, and bounded sequences. Sequence repeat-count options override converted HEIF metadata like the existing animated encoders. Legacy JPEG treats the AV1-specific lossless and bit-depth options as inapplicable and continues with its native eight-bit encoding contract. AV1 sequences now follow the existing animated-image contract: the primary item reuses the first sync sample when the root is animated, while an excluded root is encoded once as the independent primary image and the sequence begins at frame index one. Final verification remains open.
- Preserve high-bit-depth source precision through 16-bit RGB and native 10/12-bit component planes.
- [~] HEIF is registered through the default configuration module. Keep public AV1 capability claims limited to the paths covered by the encoder verification matrix until the remaining encoder work is complete.
Encoder data-flow contract:
- Resolve immutable frame and sequence output settings before allocating codec state.
- Convert each generic
ImageFrame<TPixel>once throughPixelOperations<TPixel>and the SIMD-first HEIF planar converter into native 8, 10, or 12-bit planes. Alpha is encoded as an auxiliary image when requested by the resolved output contract; it is not discarded through a source scan. - Reuse allocator-owned plane, row, block, transform, quantization, entropy, and reconstruction workspaces for the complete frame. No active path may allocate per row, block, transform, scanline, or SIMD tail.
- Analyze and encode tiles directly from those planes, retaining reconstructed reference frames only for the bounded sequence lifetime.
- Build OBU headers in bounded allocator-backed scratch and stream entropy-coded tile owners and container extents directly. Every ownership transfer is explicit, every owner is disposed exactly once, and no
ToArrayor file-sized copy crosses a layer boundary. - Iterate image frames using ImageSharp frame metadata and format-connecting metadata. Root-frame-only behavior is permitted only for an explicitly static output contract.
Encoder verification contract:
- Exercise source pixel formats independently from requested AV1 bit depth, chroma subsampling, alpha, and lossless/lossy mode.
- Run every SIMD operator through FeatureTestRunner at Vector512, Vector256, Vector128, and scalar tiers against an independent scalar oracle shaped from the same libaom revision.
- Cover discontiguous allocator buffers, constrained memory groups, cancellation, non-seekable output, multiple extents, auxiliary alpha, and bounded sequences.
- Validate produced AV1 payloads with current-main libaom and compare native planes before using ImageSharp self-decode as supplemental container coverage.
6. Build the complete AV1 frame encoder
-
[~] SIMD-first RGB-to-native-plane conversion now feeds eight-bit and high-bit-depth bordered AV1 source frames directly, preserving ImageSharp's arbitrary packed-pixel input contract without an intermediate full-frame native-plane copy.
-
[~] Auxiliary-alpha encoding now follows the same packed-pixel conversion boundary without scanning pixel contents or cloning the image. Source alpha is converted through ImageSharp's 16-bit pixel contract, deinterleaved with descending Vector512, Vector256, Vector128, and scalar traversal through the shared vector-count helpers, then scaled and rounded once by the existing native-sample writer directly into the final bordered monochrome source frame. One operation-wide allocator owner provides the packed and planar row views; there is no frame-sized alpha staging allocation or second owner. Exact 12-bit precision, physical border extension, the single 12-bytes-per-pixel row rent, and balanced return pass through the production converter. The complete 47-case frame-encoder set passes direct foreground net11 Release VSTest, and current-main
aomdecata40ed1ea9e4ecc3df58a5bccb76623f2c94ae727accepts the generated 8-, 10-, and 12-bit monochrome payloads. AVIF auxiliary item properties, references, and public activation remain open. -
[~] Forward transform families, transform workspace, and an allocation-free DC intra block boundary exist locally. For eight-bit and high-bit-depth samples, the composed boundary now follows current libaom's encoder order: predict into the reconstruction plane, subtract prediction from source, transform, quantize into separate qcoeff and dqcoeff storage, retain EOB and transform type, and inverse-transform only when EOB is nonzero so later blocks consume decoder-identical references. Prediction and subtraction retain their SIMD-first operators, independent source and reconstruction strides are preserved, and no frame-sized or per-block buffer is introduced. The block boundary consumes the real bordered encoder-plane regions and indexes their one-segment owner directly; this preserves physical row strides without a row copy and avoids the per-call enumerator allocation exposed by the initial array-only test. One reusable 61 KiB allocator owner supplies tightly packed residual, aligned transform-coefficient, dequantized-coefficient, and transform scratch spans across transform blocks; quantized coefficients write directly to the retained frame coefficient owner instead of being duplicated. A fixed 8x8 DC-intra superblock baseline now traverses the same recursive preorder and frame-edge pruning as the tile writer, gathers left references into that reusable block workspace, writes luma and chroma coefficient-owner slices in the writer's exact consumption order, and updates the caller-owned reconstruction planes for subsequent predictions. Stage-by-stage scalar-oracle, physical-border, retained-syntax, superblock-to-writer synchronization, high-bit-depth precision, and steady-state zero-allocation coverage passes 8 of 8 through direct net11 VSTest in Release. This is a legal fixed baseline, not complete partition or mode analysis.
-
[~] The production tile writer walks raster superblocks, analyzes each immediately before entropy coding, reuses one decision workspace and one block workspace, and retains decoder-identical reconstructed references across each tile. Frames exceeding AV1's 4,096-sample tile-width or 4,096-by-2,304-sample tile-area limit now select the minimum uniform tile-column and tile-row logarithms used by current libaom. Every tile begins from the same normative frame probabilities, appends its independently finalized range-coded bytes to one bounded output allocation, and records only its offset and length in the picture-state owner. Closed byte and high-bit-depth operators feed the existing superblock boundary without runtime sample-type checks. Existing byte-exact and clipped-superblock tests cover traversal and coefficient indexing; a production 4,097-sample-wide lossless case crosses the first tile boundary and checks decoded pixels on both sides. Roslynk reports zero compiler errors; runtime and current-libaom verification of this multi-tile checkpoint remain pending.
-
[~] A non-owning encoder-frame view now separates visible conversion regions from coded regions and performs complete left, top, right, bottom, and corner extension across each bordered plane. Current libaom uses 8-sample-aligned coded dimensions, a 32-sample-aligned luma stride with chroma stride derived from it, and a 64-pixel luma border for non-resized all-intra encoding. One operation-ready frame owner now rents the aligned Y, U, and V storage contiguously, exposes non-owning
Buffer2Dplane views, and returns the rent exactly once. A 4K 4:2:0 frame occupies about 13.0 MiB at 8-bit or 26.0 MiB at 10/12-bit; source and reconstruction therefore remain distinct frame owners rather than adding a full-frame copy. The corrected tests use this real ownership path and verify the exact 54 KiB 64x64 4:2:0 rent. The frame-encoder operation now instantiates matching source and reconstruction owners with ordinaryusinglifetimes and converts packed pixels directly into the source owner before extension. -
[~] Temporal delimiter, sequence header, frame header, combined-frame tile-group writing, uniform multi-tile layout, and reduced and non-reduced frame operations now exist locally. The remaining codec-tool and verification work is tracked below.
-
[~] Implement superblock and partition analysis for every permitted block size and partition. Efforts zero through eight deliberately split every in-frame node to 8x8 blocks. Effort nine performs recursive live rate-distortion selection at complete 8x8 and 16x16 nodes, while effort ten extends the same search to complete 32x32, 64x64, and 128x128 nodes. Candidate order matches current libaom:
PARTITION_NONE,PARTITION_SPLIT,PARTITION_HORZ,PARTITION_VERT, the four asymmetric partitions, thenPARTITION_HORZ_4andPARTITION_VERT_4; the two 1-to-4 partitions are excluded at 128x128 as required by current libaom. Invalid chroma geometries are excluded before evaluation. Each candidate saves and restores the exact partition, coefficient, transform, and palette neighbor edges in one aligned block-workspace owner; trials neither allocate nor copy probability state. Recursive split trials publish each selected child's decoded mode, transform, coefficient, and palette contexts before evaluating its next sibling. Coefficient contexts are published per retained transform rather than broadcasting the first transform over an entire partition leaf. Large luma and chroma leaves are evaluated as bounded-64, raster-ordered transform tiles in the existing aligned workspace, and each winning plane is copied to retained storage once. Production picture state retains the compact 8x8 mode allocation below effort nine and explicitly selects 4x4 allocation granularity when sub-8x8 partitions are enabled. Effort-dependent pruning remains. -
[~] Implement intra mode search, palette, filter intra, chroma-from-luma, and intra-block copy decisions. Live luma search now covers all 13 zero-angle base modes and all six nonzero adjustments for each of the eight directional modes. Joint spatial chroma search covers the same 61 candidates, combines both chroma planes in one rate-distortion decision, and preserves the winning shared angle adjustment. Chroma-from-luma now searches the complete signed alpha alphabet from reconstructed luma and retains its joint U/V syntax. Filter-intra now searches all five predictors after ordinary luma modes. Palette entropy, retained state, production syntax, luma and paired chroma palette selection, screen-content activation, and joint intra-block-copy mode selection exist, but their full reference decision policy and separate-encoder parity remain unverified.
-
[~] Implement inter mode search for bounded sequences, including reference selection and the decoder-supported inter tools. The sequence encoder retains the preceding reconstruction and, from effort six, searches a bounded full-pixel frame translation against that LAST_FRAME reference. Candidate discovery uses the existing SIMD-first squared-error kernels over a central analysis window, validates the winner over the complete coded luma plane, and charges its exact uncompressed-header bit count in the inter-frame rate-distortion domain. Pure translation is signaled as an identity-scale rotation/zoom model, matching current libaom's workaround for the AV1 translation-only axis defect. Each 8x8 inter-frame block first retains the complete intra candidate, then compares NEARESTMV, all three legal NEARMV dynamic-list entries, GLOBALMV, and all three legal NEWMV dynamic-list entries against it with live intra/inter, single-reference, mode, DRL, differential-vector, skip, transform, coefficient, and distortion costs. The initial NEWMV search retains the reference's cheap prediction-error stage, but every surviving mode now owns a complete transform, coefficient, skip, and distortion evaluation before mode selection; selected and candidate workspace views exchange ownership only on strict improvement. Inter trials remain in the existing shared workspace until they strictly beat the intra result, so losing trials require no backup buffer or copy. The tile writer emits the matching DRL path and normative context-selected
LAST_FRAMEreference tree instead of forcing every block through a segmentation feature. The selected DRL index reuses the filter-intra byte because those block syntax branches are mutually exclusive, preserving the existing packed state size. Packed short vectors reuse the existing picture-state owner. Effort six keeps a full-pixel fast path; effort seven refines each selected NEWMV through half- and quarter-pixel eight-tap prediction; effort eight adds the final eighth-pixel stage. The frame header advertises the matching precision, and the shared motion-vector entropy path emits fractional and high-precision symbols only when that precision permits them. Roslynk reports no compiler or scoped analyzer diagnostics, but runtime verification is pending. Additional retained reference pictures, compound prediction, and the remaining inter tools remain. -
[~] Current-libaom
av1_quantize_fp_no_qmatrixarithmetic is implemented as a closed generic forward-quantizer family with Vector512, Vector256, Vector128, and scalar paths, raster-order output, coded 64-point coefficient limits, and scan-order EOB selection. High-bit-depth paths widen before multiplying instead of applying the eight-bit coefficient clamp. Lossless blocks use the AV1 4x4 Walsh-Hadamard transform, exact lossless quantization and dequantization, four-by-four-only transform syntax, and non-skipped residual coding. Transform search and coefficient optimization remain. -
[~] Implement real rate-distortion selection and make quality and effort change work, size, and output quality. The complete luma and joint chroma candidate sets, including chroma-from-luma, filter-intra, palette, and intra-block copy, now perform live rate-distortion selection. Public quality mapping and effort tiers through exhaustive uniform luma mode/transform search are implemented. Effort nine adds exact recursive 8x8 and 16x16 partition rate-distortion selection, and effort ten extends it through 128x128; effort-dependent pruning and the remaining sequence searches remain.
-
[~] Frame effort now progressively expands the available current search: zero is DC-only, one adds every zero-angle spatial mode, two adds every legal directional adjustment, three refines the preliminary luma winner's transform type, four adds filter-intra and chroma-from-luma, and five adds adaptive palette and intra-block-copy analysis. Lower tiers do not signal unavailable sequence or frame tools, and tiers below five skip the whole-frame screen-content scan. Effort six enables
TX_MODE_SELECTand compares the winning ordinary spatial or filter-intra luma mode as one 8x8 transform against four raster-ordered 4x4 transforms; each luma palette candidate owns that size comparison from effort six onward. Effort seven searches every legal 8x8 transform type inside every ordinary spatial candidate rather than refining only the preliminary winner. Effort eight also performs the 8x8-versus-four-4x4 comparison inside every ordinary spatial and filter-intra candidate, matching current libaom's per-candidate uniform-transform ownership. Effort nine additionally searches every legal partition at complete 8x8 and 16x16 nodes in current-libaom order, and effort ten extends that recursive search through 128x128. Prediction and residual construction run once per mode and are reused across its legal transform types. A 128x128 leaf evaluates four 64x64 luma transforms and as many as sixteen 32x32 transforms per 4:4:4 chroma plane, retaining sparse state at coefficient-area offsets. Residual emission follows AV1's bounded-region order, completing Y, U, and V for each 64x64 luma region before advancing. Every 4x4 transform searches all legal types with live coefficient contexts and reconstructed intra references. The search reuses the aligned block workspace, preserves only global improvements, and performs no per-block, per-partition, or per-transform rent. Non-skipped intra-block copy writes and costs the current-libaom unsplit variable-transform root; skipped intra-block copy emits no transform-partition symbol. Effort-dependent model and transform pruning remain. Decoder-visible production cases inspect the emitted restrictions and frame state and decode the produced streams, including real effort-nine streams selecting sub-8x8 and 8x16 rectangular blocks. The complete non-HEVC HEIF/AV1 namespace passes 9,077 of 9,077 through one foreground net11 Release VSTest run. The last independently builtaomdec, from the then-currenta40ed1ea9e4ecc3df58a5bccb76623f2c94ae727snapshot, accepts the previously generated effort-eight and effort-ten payloads as well as the existing palette and intra-block-copy payloads. The affected encoder, partition, and workspace surface passes 139 of 139 through one foreground net11 Release VSTest run. The net11 Release build and Roslynk compiler and analyzer passes report zero errors. -
[~] Encoder rate accounting converts the entropy writer's live inverse cumulative distributions into current-libaom fixed-point symbol costs without allocating or duplicating probability state. Read-only luma-mode, directional-delta, filter-intra, chroma-mode, block-skip, transform-size, transform-block-skip, and complete transform-coefficient queries share the exact distributions mutated by the subsequent entropy write. Complete coefficient costing follows current libaom's optimized shape: it returns immediately for an empty transform, uses the EOB-specific base-range context, fuses magnitude, sign, base-range, and Golomb accounting into one reverse traversal, and combines repeated full base-range chunks instead of replaying each emitted symbol. Tile-lifetime level and context scratch is reused, the one-coefficient path neither clears nor initializes the forward-neighbor level map, and steady-state queries allocate nothing. Transform-size writing and costing share one subdivision-depth calculation, while shared closed symbol operations keep the writer and cost mappings for transform skip, transform type, and EOB syntax identical without forcing the estimator through the writer's slower two-pass coefficient traversal. The current-libaom fixed-point RD combiner preserves 64-bit distortion and rounds the weighted 1/512-bit rate at the required boundary. Its key-frame multiplier follows libaom's squared DC-quantizer formula and exact 10/12-bit normalization. Live final-block selection evaluates all 61 legal 8x8 luma candidates: the 13 zero-angle base modes in current-libaom order, followed by six nonzero adjustments for each directional mode. Joint chroma selection evaluates the equivalent 61 spatial candidates, combines U and V distortion plus coefficient rate, and charges one live chroma-mode and shared-angle symbol over the actual subsampled 4x4, 4x8, or 8x8 geometry. Chroma-from-luma subsamples the reconstructed luma block once into fixed-stride Q3 stack scratch, subtracts the rounded mean, evaluates all 33 signed alpha values independently for each plane with complete transform RD, and combines the cached plane results across all 1,088 valid joint pairs with one live sign cost and the conditional U/V magnitude costs. This is the allocation-free equivalent of current libaom's exhaustive 33-value path: it requires 66 evaluation transforms rather than transforming every joint pair, preserves DC-before-CfL-before-spatial tie order, and fixes the implicit chroma transform to DCT-DCT. Filter-intra follows ordinary luma candidates, searches all five predictors in syntax order, and evaluates every legal transform while reusing one prepared prediction and source residual per filter mode. Every candidate includes its live mode, angle, filter mode, alpha, and coefficient rate plus normalized pixel-domain distortion. The corrected prepared reference edges retain the common-corner prefix and width-plus-height extent required by rectangular directional prediction. A shared encoder/decoder availability calculation selects reconstructed top-right and bottom-left extensions according to tile, frame, superblock, and block reconstruction order; unavailable extensions repeat the nearest coded endpoint. Missing top or left edges retain current libaom's perpendicular-sample and bit-depth-midpoint rules. Directional prediction applies the AV1 three-degree adjustment step and reuses transform workspace for zone-three transposition before the transform overwrites it, keeping candidate evaluation allocation-free. The winning luma and chroma signed adjustments are retained in the packed final-block state consumed by the tile writer. The tile writer invokes these reusable workspace-backed selectors after mapping current neighbors and immediately before writing each block, so later decisions see reconstructed samples, coefficient contexts, and CDF updates from every preceding block. Block skip is read only after the callback has combined every coded plane. Luma and chroma candidate scratch is partitioned from the encoder's single aligned reusable block workspace; transform-size search uses that owner for four retained 4x4 transform states, local coefficient contexts, and the compact trial reconstruction needed to preserve the best result. No candidate path rents a buffer per block or per transform. Only a newly winning candidate is copied into retained frame storage. Production fixtures force every luma base predictor, both extreme adjustments in all three directional zones, available top-right and bottom-left extensions, high-bit-depth adjustment propagation, exact signed luma and chroma angle-rate terms, joint U/V decisions, packed chroma state, and 4:2:0, 4:2:2, and 4:4:4 transform geometry. The CfL fixtures derive target chroma from a pilot production encode's actual reconstructed luma through an independent scalar Q3 oracle and prove exact positive/negative alpha syntax plus zero-residual DCT-DCT reconstruction for all three subsampling geometries at 8, 10, and 12 bits. The stable fixed-DC traversal comparison uses neutral samples for which both the baseline and live search select DC with non-skip coefficient syntax, instead of relying on textured content to happen to select the baseline mode. Luma palette selection now evaluates dominant-color and one-dimensional K-means candidates for every legal size, snaps near-cache colors with the reference threshold and tie order, removes duplicate snapped colors, extends boundary maps from active samples, and performs complete transform rate-distortion search. Ordinary DC and filter-intra candidates pay the palette-disabled symbol whenever screen-content syntax is enabled. The exact net11 Release rebuild reports 1,992 test-project warnings and zero errors, all 58 intra-superblock cases pass, all 8,935 AVIF cases pass, and all 230 HEIF cases pass. Remaining mode decision work includes transform-size coverage for filter-intra and palette, broader joint mode/transform refinement, and effort-dependent pruning. Ordinary intra blocks now remain non-skipped even when all transforms are empty; inter and intra-block-copy mode selection own their distinct skip-transform RD decisions.
-
[~] The tile writer now publishes one packed coefficient context per covered 4x4 edge unit and derives luma/chroma skip plus DC-sign contexts from the complete transform edges using current-libaom units. Partition, transform, and coefficient neighbor state retains only the above and left context regions used by current libaom; the unused third top-left region, its granularity state, and its unused sentinel are removed. One picture owner packs segmentation, every tile's partition, luma, chroma, and transform edges, CDEF state, preceding quantizer, and encoded payload bounds into one clean byte allocation with typed non-owning views; together with the separately typed packed mode-information owner, the complete picture state uses two allocator rents rather than seven. Each encoded tile has independent neighbor and probability state while sharing the bounded output owner. Earlier aligned-length and balanced-return coverage exists; runtime allocation verification of the current multi-tile layout remains pending.
-
[~] Encoder mode information now uses a frame-owned integer alias grid over a packed 8-byte value allocation, matching current libaom's
mi_grid_baseandmi_allocrelationship without a managed object or reference per 4x4 entry. The visible dimensions are aligned to eight luma samples, the grid stride and allocated row count are aligned to 32 mode-information units, and optional 8x8 allocation granularity reduces the value store in both dimensions exactly as current libaom does. One clean ImageSharp byte owner contains both independently typed regions, reducing libaom's two allocation lifetimes to one without a copy. At 4K, the 4x4 layout occupies about 6.0 MiB in total; the 8x8 layout occupies about 3.0 MiB. Exact geometry, clean allocation, typed lengths, aligned mapping, untouched row padding, and exactly-once return pass 4 of 4 direct net11 VSTest cases in Release. Every coded 4x4 cell covered by square, rectangular, or clipped edge blocks maps to its owning allocation entry before context-dependent symbols are written. Packed syntax, relative neighbor lookup, full block mapping, writer traversal, entropy, and OBU coverage pass 1,947 of 1,947 direct net11 VSTest cases in Release; complete mode decision still remains. -
[~] The superblock decision and palette-map workspace uses one reusable 40.3 KiB ImageSharp allocator owner. Its aligned 8.3 KiB decision region contains 1,024 explicitly packed 8-byte final-block entries and the 341 preorder partition bytes required by a complete 128x128-through-8x8 quadtree; its remaining 32 KiB contains the fixed 128x128 luma and chroma palette maps. This combines storage held separately by libaom; exact lifetime and size reconciliation remains part of the fresh allocation audit, and fewer owners alone does not establish an improvement. Palette colors have their own current-block value and are copied only to the picture edges that later blocks can reference, so enabling palette mode does not add 50 bytes to every final-block entry. Construction and the explicit per-superblock reset initialize every syntax field, including the nonzero sentinel that disables filter-intra prediction; pooled quantizer, prediction, partition, and current-palette bytes cannot leak into the next decision pass. Roslynk reports zero compiler errors for the current one-owner refactor; runtime allocation verification remains pending.
-
[~] Finalized transform coefficients and packed EOB/type state now use raster-ordered, per-superblock plane segments matching current libaom's coefficient-pool geometry. One ImageSharp allocator owner replaces libaom's separate coefficient, EOB, and entropy-context allocations while preserving the full 1024 luma and 256-per-chroma 4x4 state capacity of a 128x128 4:2:0 superblock. The fixed 8x8 DC-intra traversal populates the owner's quantized coefficient and state slices while updating the caller-owned reconstruction plane directly, and a real tile-writer integration check proves that both sides consume identical luma and chroma areas. Complete mode decision still remains.
-
[~] Tile partition writing now follows current libaom's recursive
write_modes_sbpreorder traversal andupdate_ext_partition_contextedge updates directly. Bottom-edge blocks use the horizontal-alike partition CDF and right-edge blocks use the vertical-alike CDF; byte-exact regressions cover both paths after the previous calls were found reversed. Lossless chroma-from-luma availability now uses the subsampled plane block size shared with the decoder instead of the lossy 32x32 limit, preserving the correct UV-mode alphabet for each segment. The obsolete SVT-derived global geometry catalog and its unimplemented lookup are removed; transform geometry is derived in libaom's bounded 64x64 residual order, fixed intra transform-size symbols use the reference depth and neighbor contexts, and each derived transform size is persisted to the frame-owned mode information before the entropy snapshot and coefficient traversal consume it. Frame-edge and segmentation syntax use mode-information units, and 128x128 CDEF units use libaom's 0-to-3 indexing and first-block strength ownership. The focused transform-state regression passes 3 of 3 direct net11 VSTest cases in Release. Writer, entropy, and OBU coverage passes 1,957 of 1,957 direct net11 VSTest cases in Release, with 20 of 20 focused encoder and decoder chroma-from-luma cases. Partition and mode analysis still need to populate these retained decisions; variable inter-transform syntax remains part of later inter-frame support. -
Implement legal deblocking, CDEF, restoration, super-resolution, and film-grain signaling decisions.
-
[~] The coefficient symbol encoder allocates its bounded level and raster-context workspaces with the encoder state and reuses them for every transform and sequence sample, matching current libaom's fixed-geometry compressor lifetime instead of renting scratch during the first coded frame. Its range coder matches current libaom's 64-bit coding window, bulk big-endian byte flush, and backward carry propagation while using one byte of allocator scratch per estimated output byte instead of the former 16-bit pre-carry storage. Production tiles finalize consecutively inside one encoder-owned bounded output allocation; the OBU writer consumes non-owning tile slices synchronously before the encoder is reset or disposed, so no payload owner transfer, second rent, or full-tile copy occurs. Exact-length test callers retain the copying overload. In-memory allocation coverage proves reset-to-offset and reset-to-zero reuse the same output owner; runtime allocation verification of the complete production multi-tile and sequence paths remains pending.
-
[~] The planar conversion, DC intra prediction, residual construction, forward transform, and forward quantizer use descending SIMD dispatch: Vector512, Vector256, Vector128, then scalar. Residual construction matches current libaom's exact source-minus-prediction arithmetic for 8-bit and high-bit-depth planes, preserves independent row strides and unaligned starts, and writes directly into caller-owned signed-short storage without allocation. Candidate distortion reuses that residual workspace and widens signed 12-bit lanes before vector squaring, accumulating exact full-block SSE in 64-bit scalar storage. The composed block path delegates arithmetic to those closed operators and adds no allocation. Apply the same rule to every later hot-path family.
-
[~] Residual tests verify misaligned planes, independent source, prediction, and destination strides, SIMD remainders, untouched padding, 8-bit, 10-bit, and 12-bit precision, every operator width independently of host acceleration, the scalar fallback, and zero per-transform allocations.
-
[~] The unused coefficient-shape transform facade and its unimplemented N2, N4, and DC-only branches are removed. Finalized block encoding now follows the complete-transform path that current libaom uses before fast quantization; later rate-distortion search may add proven coefficient optimization without exposing inactive runtime throws.
-
[~] Forward-quantizer FeatureTestRunner and zero-allocation tests compare every hardware tier with an independent scan-order scalar oracle shaped from current-main libaom. Both passed direct net11 VSTest in Release.
-
[~] The combined-frame writer now completes the byte-counted uncompressed frame header before starting the optional multi-tile tile-group flag, matching current libaom's separate frame-header and tile-group writers. A non-uniform two-tile round trip verifies the explicit boundaries, both tile payloads, and complete stream consumption through direct net11 VSTest in Release.
-
[~] The first internal frame-to-OBU operation encodes 8-, 10-, and 12-bit monochrome reduced still pictures through the production tile writer and production decoder. Coefficient context initialization now stores
min(abs(level), 127), matching current libaom; the previous signed clamp converted every negative transform coefficient to zero and selected invalid nonzero-map distributions. Signed dense and sparse entropy round trips, direct level-buffer saturation coverage, and eight constant/gradient frame cases pass 52 of 52 direct net11 VSTest cases in Release. Current-mainaomdecaccepts all eight emitted payloads. After winner-mode transform refinement, their decoded-frame MD5 values ared09ea148582b9c93fa78e59426193bbc(16x16 8-bit constant),b83eedd5a84428f0120130253b30bdaa(16x16 8-bit gradient),f949f7422913e83dff07ee5e0a5087d3(8x8 8-bit constant),ae7233a94558978934469dcc4da764dd(8x8 8-bit gradient),09223b227f3abc3134d0a3ea15f70c0a(8x8 10-bit constant),539aab0e6e14bcaec271febfa8e25444(8x8 10-bit gradient),73117a8fc102e5d028f82444fc4d15ab(8x8 12-bit constant), and6936a2b62d7220dfb12f3763bb49965d(8x8 12-bit gradient). This is an independently decodable baseline, not completion evidence for chroma, alpha, options, containers, or the public encoder. -
The exact net11 Release rebuild completed at the established 1,005-warning repository baseline with zero errors. The complete HEIF/AV1 namespace passes 8,838 of 8,838 direct VSTest cases with zero failures or skips. Roslynk reports zero compiler errors and no diagnostics in the five changed C# files;
git diff --checkpasses and.gitattributesis unchanged. -
[~] The same internal frame operation now produces 4:2:0, 4:2:2, and 4:4:4 payloads at 8, 10, and 12 bits. Twenty-one color cases cover constant and spatially varying input at aligned dimensions plus odd 13x11 visible dimensions for every chroma geometry. The production decoder consumes every payload, the decoded output retains non-neutral chroma, and current-main
aomdecaccepts all 29 monochrome and color outputs. After live spatial chroma mode selection, implicit chroma-transform correction, winner-mode luma-transform refinement, exhaustive chroma-from-luma alpha selection, and filter-intra search, the odd-dimension decoded-frame MD5 values are9985f05790d2c9f5f28723ef86d5b89b(4:2:0),2ba2f1d0fcfef60394a5175553c7cb8b(4:2:2), and6aa7a2ed0dbf76ad2ec0c222585272d0(4:4:4). This proves legal current-libaom payload syntax across native plane geometries; it does not yet prove target quality or native-plane equality with an independently encoded reference. -
Spatial chroma candidates now use the implicit transform derived from the selected UV mode and active transform set, matching current libaom's
intra_mode_to_tx_typeandav1_get_tx_typebehavior. The same shared derivation is consumed by the decoder, so coefficient scan order, entropy contexts, inverse reconstruction, and encoder rate estimates cannot drift between the two paths. The previous DCT-DCT candidate transform could produce syntactically accepted streams whose non-DC chroma coefficients were interpreted under a different implicit transform. Six production mode-decision cases retain nonzero U and V coefficients and assert the selected transform state across 4:2:0, 4:2:2, and 4:4:4; fifteen exact mapping cases cover every intra mode, reduced sets, and the 32x32 DCT-only fallback. The focused contract passes 21 of 21 direct net11 VSTest cases, the complete HEIF/AV1 namespace passes 8,947 of 8,947, the exact Release rebuild remains at 1,005 warnings and zero errors, and current-mainaomdecaccepts all 29 regenerated payloads. -
[~] Luma mode selection now evaluates each of its 61 mode-and-angle candidates with the mode-derived default transform used by current libaom's fast intra path. It then refines only the winning mode across all seven transform types permitted by the 8x8 intra set in transform-enum order. This removes the fixed DCT-DCT limitation while avoiding a 61-by-7 expansion; each trial includes live transform-type and coefficient rate, reconstructed pixel-domain distortion, and the existing aligned reusable block workspace. Eighteen exact-prediction production cases prove DCT-DCT wins equal-cost ties in reference order even when the first pass used a different default, while the 72x72 textured traversal proves a non-DCT transform with nonzero coefficients reaches retained syntax. Current-main
aomdecaccepts all 29 regenerated payloads. Special-mode transform-size coverage, full partition search, and broader effort-dependent joint mode/transform search remain. -
Chroma-from-luma mode decision now reuses the decoder's SIMD-first 4:2:0, 4:2:2, and 4:4:4 reconstructed-luma preparation and prediction kernels for both byte and high-bit-depth encoder operators. The constant DC predictor for each chroma plane is computed once and its sample refills every alpha candidate, matching libaom's per-plane DC cache instead of rebuilding the same edge average 33 times. Each block uses 512 bytes of fixed stack scratch for the maximum 8-row predictor surface plus 792 bytes for complete U/V rate and distortion tables; no allocator owner, managed object, frame copy, or persistent buffer was added. Live probability costs exactly mirror current libaom's joint-sign ownership and conditional magnitude symbols. Nine production cases independently derive exact CfL targets from decoder-visible reconstructed luma at 8, 10, and 12 bits, and three entropy cases cover two nonzero signs plus each single-zero-plane form. The exact net11 Release rebuild remains at 1,005 warnings and zero errors, all 8,959 HEIF/AV1 tests pass, and current-main
aomdecata40ed1ea9e4ecc3df58a5bccb76623f2c94ae727accepts all 29 regenerated payloads. -
Filter-intra mode decision now runs after ordinary luma modes in current-libaom order, evaluates all five recursive predictors, and refines each predictor across every legal 8x8 transform in transform-enum order. Strictly-better replacement preserves ordinary-mode and filter-mode tie order. Each filter prediction and its source residual are prepared once and reused across transform candidates, avoiding repeated recursive prediction while retaining SIMD-first predictor and subtraction operators. The stack cost is 192 bytes for eight-bit samples or 256 bytes for high-bit-depth samples; no allocator owner or managed buffer was added. Fifteen production cases force every filter mode at 8, 10, and 12 bits and prove retained filter syntax, zero-residual reconstruction, and the DCT-DCT equal-cost transform tie. The decoded-frame MD5 values selected by this checkpoint are
d7d68803763b95827483f14515281d3afor the 8x8 10-bit gradient,3f7e34d44c65d7797ad26b5cd4c35bf4for the 8x8 12-bit gradient, and9985f05790d2c9f5f28723ef86d5b89b,2ba2f1d0fcfef60394a5175553c7cb8b, and6aa7a2ed0dbf76ad2ec0c222585272d0for the odd 4:2:0, 4:2:2, and 4:4:4 gradients. The exact net11 Release rebuild remains at 1,005 warnings and zero errors, 18 focused filter-intra, predictor-reference, syntax-cost, and allocation cases pass, all 8,974 HEIF/AV1 tests pass, and current-mainaomdecata40ed1ea9e4ecc3df58a5bccb76623f2c94ae727accepts all 29 regenerated payloads. -
Empty-transform block skip now compares the complete live rate of the two decoder-identical syntax choices after luma and every coded chroma plane have been selected. Current libaom forces all-intra blocks to non-skip; this encoder retains that behavior for every non-empty block and for equal-cost empty blocks, but emits block skip when its adapted context cost is strictly lower than non-skip plus all empty-transform coefficient costs. Costing and writing share the same above-and-left skip-context calculation, and the coefficient estimator returns after the transform-block-skip symbol without reading coefficient storage. This adds no allocation, copy, or persistent state. A focused adapted-CDF regression proves both outcomes through the production decision helper, the two production all-zero fixtures still prove the default real block path, the exact net11 Release rebuild remains at 1,005 warnings and zero errors, all 8,975 HEIF/AV1 tests pass, and current-main
aomdecata40ed1ea9e4ecc3df58a5bccb76623f2c94ae727accepts all 29 regenerated payloads. -
[~] Palette entropy coding now mirrors current libaom's adaptive luma-mode, chroma-mode, palette-size, and spatial color-index distributions, together with its truncated-binary uniform code used by palette colors. The complete mutable palette probability graph is created once on first palette search or write, so the current palette-disabled frame path retains zero palette allocations. Three focused regressions cover every legal 2-through-8 color alphabet and every defined mode, size, and color-index context; all 1,928 entropy cases and all 8,978 HEIF/AV1 cases pass direct net11 Release VSTest. The exact Release rebuild remains at 1,005 warnings and zero errors. This checkpoint adds the exact entropy foundation only: palette candidate generation, retained color and index storage, mode decision, map tokenization, and production syntax remain incomplete, and no generated payload changed.
-
[~] Luma and chroma palette-color coding now matches current libaom's neighbor-cache flags, sorted delta representation, wrapped V-plane deltas, strict delta-versus-raw V selection, and fixed-point color-rate model at 8, 10, and 12 bits. Encoder costing and emission use only fixed stack spans, including explicitly initialized cache-membership state, and steady-state color costing allocates zero managed bytes. The decoder consumes the same bounded color-syntax primitive after the tile reader derives its neighbor cache, removing duplicated color parsing without changing retained palette ownership. Nine focused syntax, exact palette decode, constrained-allocation, truncation, presentation, and allocation cases pass; all 1,933 entropy cases and all 8,983 HEIF/AV1 cases pass direct net11 Release VSTest. The exact Release rebuild remains at 1,005 warnings and zero errors. Retained encoder palette colors, neighbor caches, color-index maps, candidate generation, and production palette selection remain incomplete, and the compact 8-byte frame mode entries were not enlarged.
-
[~] Palette color-index map coding now shares the exact current-libaom neighbor weights, stable color ordering, five context classes, first-index uniform code, and diagonal wavefront between encoder costing, encoder writing, and decoder parsing. The decoder's stack-allocated context scores are explicitly cleared before accumulation, removing an invalid dependency on uninitialized stack contents. Costing and writing use a closed generic operation while the shared driver owns traversal and context derivation, so the semantic operations remain independent of map layout and tail handling. The path adds no retained state or per-call managed allocation. Its allocation regression now runs one complete unmeasured hot-path window before measuring an independent 1,000-call steady-state window, so tiered-runtime transitions cannot make the full parallel suite report a one-time allocation as a recurring operation cost. Twelve focused map, exact palette decode, padding, trailing-bit, and allocation cases pass; all 1,941 entropy cases and all 8,991 HEIF/AV1 cases pass direct net11 Release VSTest. The exact Release rebuild remains at 1,005 warnings and zero errors. Production payloads remain unchanged because palette selection is still disabled; retained colors, neighbor caches, index-map storage, candidate generation, and production palette mode decision remain incomplete.
-
[~] Retained encoder palette state and production palette writing now mirror current libaom's 50-byte palette-mode contents, separate luma and shared-chroma sizes, three eight-color planes, above-and-left sorted cache, 64-sample above-cache boundary, mode contexts, palette colors, color-index maps, and syntax order. The current block keeps one inline value in the reusable superblock workspace; only the 4x4-granularity top and left picture edges retain copies for later blocks. For a 3840x2160 tile these edges occupy about 73.4 KiB instead of about 6.2 MiB for a 50-byte palette value attached to every 8x8 mode allocation. Luma and chroma index maps occupy a fixed 32 KiB region of the single 40.3 KiB superblock-workspace owner. That owner is allocated with encoder state, matching libaom's compressor-state lifetime while removing libaom's separate palette allocation and cleanup path. The compact final-block decision region remains about 8.3 KiB. The writer caps map traversal to the coded plane count, writes maps before transform syntax, and publishes palette edges only after the current block has consumed preceding contexts. The previous eight focused size, alignment, ownership, cache-boundary, round-trip, map-consumption, and edge-publication cases passed with all 114 palette cases, all 1,942 entropy cases, and all 8,996 HEIF/AV1 cases through direct foreground net11 Release VSTest. The current one-owner refactor has zero Roslynk compiler errors; runtime verification remains pending. The current source reference is official libaom main at
d565eec60f084421fa34fc0534b760c6452b6a6c. -
[~] Luma palette clustering now follows current libaom's one-dimensional search primitive exactly: equal-interval midpoint initialization, first-color tie order, rounded centroid means, deterministic empty-cluster replacement, the 50-iteration limit, and retention of the preceding state when distortion increases. Nearest-color assignment dispatches Vector512, Vector256, Vector128, then scalar through ImageSharp's shared vector-count helpers. Wider dispatch alone does not establish an end-to-end performance improvement. The primitive uses only bounded stack scratch and introduces no allocator rent, managed array, or per-row copy. Three independent tests cover exact centroid convergence, initialization order, 12-bit nearest-color distortion, destination bounds, and every hardware-intrinsic tier. The complete AVIF set passes 8,930 of 8,930 cases and the HEIF set passes 230 of 230 cases through direct foreground net11 Release VSTest. The exact net11 Release rebuild reports 1,050 solution warnings and zero errors, and Roslynk reports zero compiler errors. Candidate enumeration, palette-cache snapping, transform RD selection, and production activation remain in the open luma-palette checkpoint.
-
[~] Live luma palette selection now follows current libaom's dominant-color and one-dimensional K-means candidate families, cache-bias threshold, sorted duplicate removal, active-edge map extension, and strict winner tie order. It evaluates both candidate families at every legal 2-through-8 size without the reference's speed-dependent pruning, then evaluates every legal transform. This is a controller deviation; evaluating more candidates does not establish improved quality, speed, or reference parity. Candidate storage remains bounded stack memory; the reusable maps come from the fixed encoder-lifetime superblock workspace, so palette search cannot introduce a first-use allocation. A production tile test proves that full 8x8 and clipped 5x3 blocks at 8 and 12 bits select exact colors and indices, extend the visible edges through coded padding, reconstruct every sample without coefficients, and emit a nonempty tile. The complete 57-case intra-superblock set, 8,931-case AVIF set, and 230-case HEIF set pass direct foreground net11 Release VSTest. The exact Release test-project build reports 1,992 baseline warnings and zero errors; Roslynk reports zero compiler errors and no touched-file analyzer warnings. Production frame activation remains gated until chroma palette mode and its rate accounting are complete.
-
[~] Paired chroma palette clustering now preserves current libaom's squared two-component distance, first-centroid tie order, independently rounded U/V means, paired deterministic empty-cluster replacement, preceding-state retention on increased distortion, and 50-iteration limit. The source planes remain separate, with Vector512, Vector256, Vector128, then scalar dispatch through ImageSharp's shared vector-count helpers. An end-to-end improvement over the native implementation has not been established. Three independent tests cover exact paired convergence, midpoint initialization, 12-bit distance and index parity, untouched destination bounds, and every intrinsic tier. The exact Release test-project build reports 1,992 baseline warnings and zero errors; the focused three-case set, complete 8,934-case AVIF set, and complete 230-case HEIF set pass direct foreground net11 Release VSTest. Roslynk reports zero compiler errors and no touched-file analyzer warnings. Candidate integration and production activation remain in the open chroma-palette checkpoint.
-
[~] Live paired chroma palette selection now follows current libaom's complete 2-through-8 color-size search, U-plane neighbor-cache snapping, stable U-ordered color pairs, shared U/V index map, implicit DCT-DCT transform, and strict rate-distortion winner replacement. It omits the reference's early header-cost pruning, keeps planar U/V source data separate, and reuses the existing prediction, residual, transform, quantization, and reconstruction operators. Omitting pruning is an unresolved decision-policy deviation, not an established improvement. The production tile regression proves both palette-mode probability branches, exact paired colors and indices, coefficient-free reconstruction, and nonempty syntax. The complete 58-case intra-superblock set, 8,935-case AVIF set, and 230-case HEIF set pass direct foreground net11 Release VSTest. The exact Release test-project build reports 1,992 baseline warnings and zero errors; Roslynk reports zero compiler errors and no touched-file analyzer warnings. Production frame activation remains the next checkpoint.
-
Production screen-content activation now matches current libaom's default good-quality detector: it scans only complete 16x16 luma blocks, normalizes palette samples to eight bits, admits 2-through-4-color blocks, and uses the reference's strict greater-than-ten-percent frame-area threshold. The same pass accumulates centered sums and squared sums at native precision, applies libaom's exact 10-bit and 12-bit variance rounding, and enables intra-block copy only when positive rounded per-pixel variance exceeds its strict one-twelfth frame-area threshold. A 256-bit stack bitset and fifth-color early exit replace libaom's larger per-block histogram without a second source scan or allocation. The adaptive sequence flag remains enabled and both frame flags are fixed before picture-state allocation. Focused regressions prove strict palette-threshold equality, high-bit-depth normalization, the exact variance rounding boundary, five-color rejection, emitted frame-header activation, actual production IBC selection, and production decode. The exact Release test-project build reports 1,992 baseline warnings and zero errors; all 9,242 non-HEVC HEIF/AV1 cases pass direct foreground net11 Release VSTest. Current-main
aomdecata40ed1ea9e4ecc3df58a5bccb76623f2c94ae727accepts all 31 payloads regenerated by the current test tree, including an actual IBC-coded 328x16 stream with decoded MD5677435e5af39c930af1178f91c34af6a. Roslynk reports zero compiler errors and no touched-file analyzer warnings. -
[~] Intra-block-copy rate accounting now uses the live frame-local flag and displacement-vector distributions without copying or adapting either context during candidate measurement. Displacement-vector costing and writing share one closed symbol operation over the exact current-libaom joint, sign, magnitude-class, class-zero, and integer-offset syntax; final mode evaluation applies libaom's 120/128 displacement-rate weight with nearest-integer rounding. Independent fixed costs cover all four joint states, both signs, class zero, and large offset classes before adaptive writes, followed by an encoder/decoder round trip through the same sequence. Encoder and decoder reference-vector derivation now share the exact eight-candidate spatial scan, independent nearest and outer-region ranking, top-right partition geometry, clamping, and tile-relative fallback. Selected vectors use a naturally aligned pair of signed 16-bit components packed into the existing picture-state owner only when intra-block copy is permitted; a 3840x2160 frame retains 130,560 vectors in 510 KiB while leaving the compact 8-byte mode allocation unchanged. The tile writer derives the same reference and emits the retained vector without another allocation or copy. Coefficient costing and writing now select the inter transform sets and frame-local probability tables required by intra-block copy; independent tests verify every legal symbol against the exact default inter distribution and round-trip full and reduced sets from 4x4 through 32x32. Legal 8x8 hash discovery now indexes every visible source origin, including unaligned origins, in libaom's coarse-to-fine insertion order with the same 256-candidate bucket cap. A separable rolling hash fills one packed picture-lifetime workspace before reconstruction, then reuses that workspace for integer candidate links; exact wide or SIMD block comparison rejects hash collisions, and SIMD variance uses libaom's eight-bit normalization at 8, 10, and 12 bits. Power-of-two bucket arrays scale down with small images and stop at the reference's 16-bit limit, avoiding libaom's fixed six-size pointer table; the 3840x2160 search index occupies about 32.2 MiB and introduces no additional owner or frame copy. Above and left search rectangles, integer displacement legality, strict tie order, and live raw displacement rate follow current libaom. Motion-candidate ranking uses libaom's undiscounted probability cost and exact variance-domain error-per-bit scaling, separately from the later 120/128 final-mode discount. The allocation-free full-pixel core now follows current libaom's NSTEP search: it clamps the spatial reference to each legal region, traverses the fixed 15-stage radii and site order, skips equivalent centered 210-pixel stages, repeats progressively shorter paths, and compares their winners in the normalized variance domain. Paths above the speed-zero screen-content threshold continue through libaom's 256-pixel, one-pixel-step exhaustive mesh. Four adjacent byte or high-bit-depth candidates share each SIMD source load, strict row-major tie ordering is retained, and the final legal tail column remains searchable where libaom's current four-wide remainder loop omits it. Byte and high-bit-depth operators compute each 8x8 absolute difference with Vector128 before scalar fallback; high-bit-depth SAD remains in its native sample scale while its quantizer-derived rate multiplier uses libaom's normalized AC step. Production mode decision now derives the same spatial displacement reference used by the writer, deduplicates hash and full-pixel finalists in search order, and evaluates every surviving vector through complete luma and chroma transform RD. This differs from libaom's preliminary-error pruning by permitting a hash and pixel finalist from the same search region to compete using final syntax and reconstruction costs. It is an unresolved controller deviation, with no established quality or performance improvement. Prediction is prepared once per plane and vector, including integer or half-sample chroma phase, then reused across every legal inter transform without an allocator rent or frame copy. The joint comparison includes the live intra-block-copy flag, discounted displacement rate, skip flag, coefficient syntax, and normalized Y/U/V distortion; an empty transform alternative can win only when its complete skip cost is strictly lower, while conventional intra and earlier vectors retain tie precedence. Winning reconstruction, coefficients, transform state, DC modes, cleared palette/filter/CfL state, and displacement are copied once into the existing retained stores. Production regressions force the path at 8 and 12 bits and force 4:2:0 horizontal half-sample chroma with an unaligned reference. The former bulk local workspace occupied 2.75 KiB for byte samples or 3.375 KiB for high-bit-depth samples. Prediction, candidate, winning reconstruction, residual, and coefficient scratch now occupy one naturally aligned 3.125 KiB extension of the existing frame-reused block-workspace owner, matching libaom's reusable macroblock-scratch lifetime without adding an allocation; only the 128-byte reference, weight, and finalist arrays remain on the stack. The net11 Release solution build reports zero errors; all 2,082 focused transform, entropy, intra-block-copy, intra-superblock, and frame-encoder cases and all 9,242 non-HEVC HEIF/AV1 cases pass through direct foreground VSTest, with tiered compilation disabled only for the full allocation-sensitive suite. Adaptive production activation is complete, and the emitted frame flag remains authoritative for the complete frame rather than being invalidated after tile coding.
-
The expanded checkpoint exposed a pre-existing transform-block test that asserted uninitialized pooled padding was zero. The test now initializes the complete physical luma plane with a sentinel and proves the block operation leaves both adjacent padding samples unchanged. The exact net11 Release rebuild remains at 1,005 baseline warnings and zero errors, the focused allocator-order set passes 30 of 30 cases, and the complete HEIF/AV1 namespace passes 8,859 of 8,859 direct VSTest cases with zero failures or skips.
-
Combined-frame OBU output now counts the byte-aligned frame and tile-group headers, non-final tile-size fields, and owned tile payloads before emitting the OBU size. It retains only the small allocator-owned header scratch and writes each entropy-coded tile span directly from its detached owner, removing the second file-sized allocator rent and complete-payload copy. A 64 KiB regression proves exactly one sub-payload-sized byte rent with a balanced return and verifies the exact streamed tile tail; the existing two-tile round trip proves size-prefix and ordering parity. The focused writer and production-frame set passes 32 of 32 direct net11 VSTest cases, current-main
aomdecaccepts all 29 generated native-format payloads, and the complete HEIF/AV1 namespace passes 8,860 of 8,860 cases with zero failures or skips. -
The earlier fixed-block skip checkpoint was not equivalent to libaom's ordinary-intra policy. Its all-zero-EOB conjunction produced valid streams, but ordinary intra retains non-skip syntax in the reference. The takeover correction above aligns both fixed-DC traversal and live mode selection. Earlier decoder acceptance, unchanged frame hashes, and one-byte size reductions did not establish encoder-policy parity; the earlier 8,862-case result is historical evidence only.
-
Operation-wide allocation tracking now exercises a real 64x64 12-bit 4:4:4 frame through packed-pixel conversion, both native frame owners, picture and coefficient state, reusable block workspaces, entropy coding, OBU framing, and a non-seekable destination. It proves exactly one 60 KiB tile-output reservation from current libaom's all-intra 2.5x rule and balanced exactly-once returns for every tracked allocation before the operation completes. The focused ownership case passes 1 of 1 and the complete HEIF/AV1 namespace passes 8,863 of 8,863 direct net11 VSTest cases with zero failures or skips.
-
HEIF box offsets are now counted from the start of the encoded file instead of reading
Stream.Position. This preserves ISO BMFF file-relativeilocoffsets when the destination begins at a nonzero position and permits non-seekable output. Decoder item extents and image-sequence chunk offsets now resolve from that same file origin rather than the backing stream origin. Real legacy-JPEG HEIF round trips cover non-seekable output and a prefixed destination, while current-position AV1 decode covers both a still item and a five-frame sequence. All 96 encoder/decoder cases and all 38 sequence-parser cases pass direct net11 Release VSTest; the Release build remains at the established 1,005-warning baseline with zero errors. -
[~] Current-libaom source comparison now drives uniform luma transform ownership at each effort boundary. Effort six retains the cheaper winner-only size decision for ordinary spatial and filter-intra modes, while every palette candidate already owns its size decision. Effort seven evaluates every legal 8x8 transform type for every ordinary spatial candidate. Effort eight and above make transform size part of every ordinary spatial and filter-intra candidate's rate-distortion result, so an 8x8-only preliminary comparison cannot discard the mode that wins with four 4x4 transforms. Prediction and subtraction are prepared once per mode and reused across transform types, matching the reference separation between prediction and transform search. Each 4x4 transform searches every legal type with coefficient contexts derived from retained transform edges and preceding trial blocks, while reconstructed top-right and bottom-left references follow production coding order. Palette prediction uses non-owning subregions of the retained color map, and filter-intra rebuilds each recursive prediction from reconstructed edges. The existing aligned block-workspace owner retains prediction, residual, coefficients, contexts, compact reconstruction, and four final states; no allocator rent, managed array, best-candidate re-transform, or full-block intermediate copy was added. Dense decision points now document scratch lifetime, enumeration tie order, global-winner publication, raster reconstruction dependencies, and the deliberate lower-effort shortcut. The packed encoder transform edges initialize to 64, matching libaom and the ImageSharp decoder before a coded neighbor publishes its size, and variable transform syntax remains gated to blocks larger than 4x4. The focused Release verification passes 13 of 13 cases across efforts zero through eight and ten, palette split selection, and transform-size selection. The complete non-HEVC HEIF/AV1 namespace passed 9,301 of 9,301 cases at that checkpoint. The
aomdecbuilt from the then-currenta40ed1ea9e4ecc3df58a5bccb76623f2c94ae727snapshot accepts the generated effort-eight and effort-ten streams. Partition search and effort-dependent pruning remain. -
Intra-block-copy transform search now prepares motion compensation and subtraction once per plane, alternates the existing candidate and selected work buffers whenever a transform improves, and performs at most one final normalization copy into the caller-owned selected span. This matches current libaom's pointer-swap ownership without adding an allocation or a third reconstruction buffer. Inline documentation now records the scratch lifetime, strict transform tie order, skip-rate replacement, unsplit transform-root syntax, joint-plane winner retention, and final publication boundary. The focused Release encoder and intra-block-copy set passes 17 of 17 cases, the complete non-HEVC HEIF/AV1 namespace passes 9,301 of 9,301 cases with zero failures or skips, and current-main
aomdecaccepts the regenerated effort-five and effort-six intra-block-copy streams. -
Uniform four-by-four luma transform search now uses two compact reconstruction and coefficient views already available in the aligned mode-decision workspace. Legal transform trials write into the non-winning view and exchange span ownership only on strict rate-distortion improvement; the selected coefficients and strided reconstruction mosaic are published once after the type search so the next raster transform sees the required decoded edge. This removes reconstruction and coefficient copies on every improving transform without adding storage, changing tie order, or repeating a transform. All four representative palette, filter-intra, and effort-eight output hashes are unchanged, the focused Release set passes 13 of 13 cases, and current-main
aomdecaccepts every checked stream. -
Candidate distortion now follows the reference separation between immutable source residuals and reconstructed-pixel error. The forward-transform boundary accepts a read-only residual, so all transform types for one prediction reuse that block directly instead of copying it into transform scratch before every trial. The shared residual API now measures strided source-versus-reconstruction squared error with documented Vector512, Vector256, Vector128, and scalar traversal through ImageSharp's vector-count helpers, eliminating the former residual destination write and second reduction pass. The same path covers ordinary intra, filter intra, palette, chroma-from-luma, split transforms, and intra-block copy at 8, 10, and 12 bits without an allocation. Independent scalar, stride, tail, intrinsic-tier, and zero-allocation coverage passes with the 140-case focused encoder set; the complete non-HEVC HEIF/AV1 namespace passes 9,301 of 9,301. Representative palette, filter-intra, and effort-eight output hashes remain byte-identical, and current-main
aomdecaccepts every checked stream. -
Lossless still-image coding now follows current libaom's qindex-zero path without introducing a per-block allocation or a second frame buffer. The forward 4x4 Walsh-Hadamard transform and its transpose into the entropy pipeline's row-major coefficient order are allocation-free at Vector512, Vector256, Vector128, and scalar tiers; lossless quantization reconstructs the original transform coefficient exactly. The mode decision fixes lossless transforms to DCT-DCT syntax and four-by-four blocks, disables transform skip for nonzero residuals, and excludes the fixed-eight-by-eight intra-block-copy search that cannot represent the required lossless transform grid. The frame coefficient owner reserves the exact worst-case 2,048 transform states needed by both 128x128 4:4:4 chroma planes without adding an allocation. The color configuration derives its plane count from the monochrome flag, so high-bit-depth direct-frame writers and readers cannot retain contradictory mutable state. Public 8-bit, 10-bit, and 12-bit color and auxiliary-alpha round trips are pixel exact on their native sample lattices; direct 10-bit and 12-bit 4:4:4 frame round trips are also exact at native-plane precision. The Release build completes with zero errors, Roslynk reports zero compiler errors, and the complete non-HEVC HEIF/AV1 namespace passes 9,081 of 9,081 through one foreground VSTest run. An independently built generic
aomdecfrom current official libaommainatd565eec60f084421fa34fc0534b760c6452b6a6caccepts all 59 current payloads, including the public color and auxiliary-alpha lossless streams at every supported precision.
7. Write complete AVIF output
- [~] The encoder-side AV1 codec configuration is now derived directly from the encoded sequence header and writes the fixed four-byte
av1Crecord with emptyconfigOBUs. The image payload retains the required sequence header, so the property introduces no sequence-header allocation, retention, or copy. Four production-header cases cover main, high, and professional profiles; 8-, 10-, and 12-bit precision; monochrome, 4:2:0, 4:2:2, and 4:4:4 sampling; exact fixed bytes; decoder reparsing; and header/property equivalence through direct net11 Release VSTest. Property-container emission and public AVIF activation remain open. - [~] AV1 image properties now write
ispe,pixi,av1C,colr, andauxCin current AVIF item order. Onlyav1Cis essential; color and alpha items retain independent property sets and the registered alpha auxiliary type. The property container reacquires its span after nested expansion before patchingipco, removing the prior stale-buffer write, and selects compact or 15-bitipmaindices from the property count rather than the unrelated item count. A forced-growth color-plus-alpha case validates every property payload and association byte; a separate 43-item, 129-property case proves indices 127 through 129 and the extended essential bit. Both pass direct foreground net11 Release VSTest. Complete AVIF assembly remains open. - [~] Explicit public AV1 encoding now writes a still-image AVIF with
avifmajor brand, compatibleavif,mif1, andmiafbrands, one primary color item, an optional alpha auxiliary item,auxlfrom alpha to color, independent item properties, absolute version-oneilocextents, and a sharedmdat. Quality uses current libaom's quantizer-to-qindex mapping with public quality 100 deliberately clamped from lossless qindex 0 to qindex 4. Effort controls the implemented search stages, and the resolved value is required explicitly by every internal frame, tile, and mode-decision operation rather than repeated as optional defaults. Encoder options take precedence over source metadata for 8-, 10-, and 12-bit monochrome, 4:2:0, 4:2:2, and 4:4:4 output. Alpha derives from the source pixel type without scanning pixels, and incompatible identity-matrix metadata is normalized without mutating the source image. - [~] The production path writes color and alpha payloads sequentially through allocator-backed chunked storage, supports non-seekable and prefixed destinations, and does not materialize a complete file or payload copy. Uniform encoder-side
pixidepth is written directly without allocating per-item channel-depth arrays; decoder-side non-uniform channel depths remain supported. The Release test project builds with zero errors, all 39 encoder cases pass, the complete non-HEVC HEIF namespace passes 9,277 of 9,277, and current official libaom accepts all 47 generated payloads. - Still-image AVIF metadata preservation now writes an unrestricted ICC
colr/profproperty before the independentcolr/nclxproperty, Exif and XMP as separatemdatitems, and onecdscrelationship from each metadata item to the primary color item. Exif stores the exact big-endian TIFF-header offset required by the HEIF item syntax; XMP uses themimeitem type andapplication/rdf+xmlcontent type. Existing ICC and XMP storage is read synchronously and copied once into final encoder storage rather than cloned into an intermediate array.SkipMetadatasuppresses all three profile types while retaining the CICP values required to describe the encoded planes. The same option now reaches legacy JPEG payloads, whose encoder no longer writes application profiles or comments when metadata is disabled. - Exact container tests verify every emitted item declaration, name, MIME content type,
cdscrelationship, Exif offset and payload, XMP payload, ICC/CICP property order, compact association byte, propertyless metadata exclusion, decoded profile value, and bothSkipMetadatabranches. The final HEIF encoder set passes 44 of 44 and the complete JPEG encoder set passes 257 of 257 through direct foreground net11 Release VSTest. The complete non-HEVC HEIF namespace passes 9,282 of 9,282 with no failure, crash, or detached test host, and current official libaom accepts all 47 current generated AV1 payloads. - A code-wide production HEIF/AV1 stack-storage audit, excluding HEVC, removed every block-sized, variable-length, or repeatedly nested scratch buffer. Spatial luma and chroma, filter-intra, chroma-from-luma, luma and chroma palette selection, and K-means iteration now use typed views over 642 signed-integer elements, about 2.51 KiB, at the start of the shared inter-prediction region. Those searches are sequential for one block, so the block-workspace owner does not grow and no rent, copy, or additional lifetime is introduced. CDEF directions, variances, and its 64-entry block list now append 1 KiB to the existing bounded operation owner instead of occupying hidden inline or explicit stack arrays. No remaining
stackallocdepends on block dimensions, sample count, or runtime length; the largest remaining individual span is 128 bytes, and the remaining sites are fixed syntax, SIMD-lane, filter-tap, plane-metadata, or small candidate storage. The exact-owner test now proves the mode, palette, and reference-prediction views share one allocation. Roslynk reports zero compiler errors and no diagnostics in the changed files, the Release test-project build completes with the established 1,992 warnings and zero errors, 81 of 81 focused cases pass, and the complete non-HEVC HEIF/AV1 namespace passes 9,282 of 9,282 through one foreground net11 VSTest run. - [~] Bounded public image-sequence output now emits an
avismovie with version-one movie, track, and media headers; AV1 visual sample entries; exact run-length-compressed timing; per-sample sizes; 64-bit chunk offsets; and an explicit sync-sample table. The file type includes the requiredmiafcompatibility brand, and every sequence now has the MIAF primary image item emitted by current libavif: normal sequences share the first sync-sample extent without another encode or copy, while separate-root sequences retain the still root as the primary image and begin timed samples at frame index one. The decoder allocates one finalImage<TPixel>: the root is either the first timed sample or the separately decoded primary item, and each visible timed sample is decoded directly into a frame owned by that image. Exact quarter-turn presentation uses one frame-sized reusable pre-rotation buffer rather than a second image or a separately built frame collection. Color and optional auxiliary alpha use independently configured AV1 tracks linked byauxl. Lossless samples remain independently decodable key pictures and repeat the sequence header required for random access. Lossy continuation samples use LAST_FRAME inter prediction through the existing SIMD translational predictor; one track-scoped encoder session reuses its source allocation, packed-to-planar row storage and color converter, frame-sized coefficient storage, fixed-geometry picture and frame-header syntax state, tile/superblock/entropy cursors, block arithmetic workspace, complete probability graph, bounded tile-output owner, and OBU-header owner, and swaps two complete reconstruction buffers so the preceding decoded frame becomes the next reference without a plane copy. Sequence samples signalstill_picture=0and use the complete non-reduced sequence and frame-header prefixes required for a multi-frame coded sequence. Frame payloads are written once into contiguous allocator-backed chunks per track. One compact managed table retains only offset, length, and duration for both tracks, and the boundedmoovowner is patched once after its final size is known, so prefixed and non-seekable destinations require neither seeking nor a file-sized copy. The media timescale uses the exact representable least common multiple of animated frame-delay denominators and a documented microsecond fallback; zero delays become the smallest legal positive duration. Public lossless color-and-alpha round trips preserve all frames, distinct 24, 25, and 30 fps delays, finite or infinite repetition, ICC, Exif, and XMP metadata, while a separate case proves prefixed non-seekable output. Per-tile CDEF preset, preceding-quantizer state, and payload bounds now occupy one aligned region in the reusable allocator-owned picture buffer rather than separate managed arrays for every frame. The last verified net11 Release checkpoint completed with the established 1,005 warnings and zero errors, all 48 HEIF encoder cases passed, and the complete non-HEVC HEIF/AV1 namespace passed 9,317 of 9,317 through foreground VSTest. Current official libaommainatd565eec60f084421fa34fc0534b760c6452b6a6caccepted all 66 raw AV1 payloads regenerated by that suite. Verification of the current primary-item, separate-root, non-reduced sequence-header, retained-reference continuation, grid implementation, block-local motion search, conversion-row, probability, picture, frame-header, tile-cursor, tile-output, and OBU-header reuse, and multi-tile output is pending. Additional reference roles and compound prediction remain open. - [~] Oversized still-image encoding now writes a derived AVIF grid when either source dimension exceeds the AV1 frame-header limit. Cells are encoded row-major from source rectangles without cropping to temporary images. Every coded cell uses the same at-most-65,536-sample extent for current-reader interoperability, with edge replication supplying the AVIF minimum 64-sample dimension and any final-row or final-column crop. Color and optional alpha grids use hidden AV1 items, ordered
dimgreferences, one shared property set per plane, and independent descriptor payloads. The item-property length calculation counts every reusedipmaassociation while retaining oneipcoproperty definition. Explicit subsampled output rejects grid dimensions that MIAF cannot represent; an unspecified sampling choice promotes that exceptional odd-dimension grid to 4:4:4. Exact descriptor, hidden-flag, reference-order, property-reuse, cell-padding, and public round-trip coverage is present, and Roslynk reports zero compiler errors. Runtime verification remains pending. - [~] Write the correct AVIF file type, item information, locations, references, properties, AV1 configuration, dimensions, color, alpha, metadata, and media data.
- [~] Support single images, alpha auxiliary images, grids, multiple extents, and bounded image sequences in the final public scope.
- Preserve ICC, Exif, and XMP according to encoder options.
- [~] Write CICP, range, chroma position, bit depth, and subsampling values that match the encoded planes.
- Encode the current ImageSharp pixel lattice without inventing HEIF clean-aperture, rotation, or mirror properties. The HEIF decoder materializes those container transforms before returning an image, while ImageSharp encoders consistently preserve explicit Exif metadata without implicitly changing the pixels.
- [~] Stream output through allocator-backed chunked storage without file-sized copies or ToArray materialization.
Encoder exit gate:
- Current-main libaom accepts the payloads regenerated from the final tree.
- Reverify lossless output at public pixel and native-plane precision for 8-, 10-, and 12-bit output.
- Separately encoded lossy outputs differ by no more than one unit at every decoded output sample with reconciled settings; report maxima and counts exceeding one.
- Record equivalent end-to-end absolute timing, output size, quality, and allocation evidence after the source audit.
- 8, 10, and 12-bit monochrome, 4:2:0, 4:2:2, and 4:4:4 outputs pass.
- Alpha, grids, metadata, color profiles, transforms, and bounded sequences pass.
- ImageSharp decode of its own output is supplemental coverage only, never the sole oracle.
- Verify every exposed encoding combination through established ImageSharp conversion behavior.
- Focused Release and FeatureTestRunner verification passes with exact recorded evidence.
Architecture rules
- Follow the JPEG color-converter operator architecture exactly.
- Each distinct prediction traversal owns a family-named predictor type.
- The family .Operator.cs file defines the nested static operator contract.
- Each semantic readonly struct belongs to that owner and implements concrete scalar, Vector128, Vector256, and Vector512 arithmetic for the shared traversal.
- Do not place a distinct predictor beneath a broad Av1IntraPredictor or Av1InterPredictor.
- Do not create semantic forwarding wrappers, top-level operator types, hardware-width-named operator types, CRTP contracts, or one file containing unrelated semantic operators.
- Forward transforms belong to Av1ForwardTransformer and its semantic operator files.
- Inverse axis transforms belong to Av1Inverse2dTransformer and its semantic operator files.
- Reconstruction output operators belong to Av1InverseTransformer.
- Shared lane primitives belong only in explicitly named Operations types.
- Dispatch from widest to narrowest supported SIMD width, then execute one scalar tail.
- Keep codec execution sequential. Do not add parallel execution inside the codec.
- Do not allocate per row, block, transform, scanline, or SIMD tail.
- Use ImageSharp allocators and pools. Do not use ToArray to cross an ownership boundary.
- On internal types, use public members when other types consume them; reserve private members for type-local behavior.
- Use established ImageSharp test data, allocator tracking, FeatureTestRunner, and reference-image comparison APIs. Do not build custom substitutes.
- Public XML documentation describes observable behavior only.
- Inline comments explain the current-libaom numerical rule, ownership boundary, edge extension, entropy ordering, or SIMD shape at technically complex points.
- Every multiline statement or declaration is followed by vertical whitespace.
- Do not edit .gitattributes directly.
- Do not install or download tools without explicit permission.
Final verification matrix
- Release source build: net10.0, zero errors.
- Release source build: net11.0, zero errors.
- Scoped semantic inspection: zero compiler errors attributable to this work.
- Focused decoder syntax, reconstruction, ownership, presentation, and malformed-input tests.
- Focused encoder syntax, payload, container, precision, ownership, and option tests.
- FeatureTestRunner coverage for normal, narrower SIMD tiers, and scalar fallback.
- Constrained multi-group allocator coverage with balanced exactly-once returns.
- Exact native-plane comparisons against current-main libaom.
- Established final-presentation comparisons at the target pixel precision.
- Scoped StyleCop and vertical-whitespace inspection.
- No stale unsupported capability claims or removed-code references.
- No restore-source failures, background test hosts, detached processes, or crash-report popups.
- .gitattributes unchanged.
- git diff --check clean.
- Documentation records exact commands, counts, current-main reference revision evidence, and results.
- Commit only after the relevant checkpoint is genuinely complete.
- Do not push.