135 KiB
AVIF and AV1 implementation plan
Goal
Complete a production-quality, fully managed AV1 codec and its bounded AVIF/HEIF image container integration for ImageSharp. The finished work must decode and encode still images and bounded image sequences, preserve source precision, use ImageSharp memory ownership, and provide SIMD-first hot paths with one behaviorally identical scalar fallback.
This plan is the authoritative delivery checklist. A source file, unit test, build, self-roundtrip, or local implementation is not completion evidence by itself.
Source authority
- AV1 codec syntax, tables, fixed-point arithmetic, prediction, transforms, entropy behavior, filters, encoder decisions, and lifecycle behavior must be ported and checked only against the current
mainbranch of the official libaom checkout atD:\GitHub\AOMediaCodec\aom. - Libaom is the sole external codec implementation source. Do not use HM, libheif, FFmpeg, GPAC, SVT-AV1, dav1d, libgav1, or any other codec implementation as an algorithm, arithmetic, output, or architecture reference.
- Existing ImageSharp and JPEG code is authoritative only for ImageSharp architecture, allocator ownership, SIMD dispatch, pixel conversion, and test API patterns. It is not an alternate AV1 algorithm source.
- Production code must not load, invoke, install, or fall back to a native codec.
- Existing independent container files may be used only as interoperability inputs. Native AV1 expected output must be generated by the current libaom
maincheckout, and no independent decoder output may substitute for it.
Reference checkout evidence on 2026-08-31:
D:\GitHub\AOMediaCodec\aomis clean. Its checked-outHEADis441c439b9916474cac15d2822af47a9ad70674a8, while the refreshedorigin/mainisa40ed1ea9e4ecc3df58a5bccb76623f2c94ae727.- Current encoder verification uses the exact
origin/maintree exported toD:\GitHub\ynse01\aom-main-a40and built inD:\GitHub\ynse01\aom-main-a40-build. The resultingaomdecidentifies itself as version 3.15.0. This records the tree audited on that date; it is not a pin and must not prevent later work from updating to the then-currentmain.
Status notation
- Verified: the current behavior has exact evidence from the current libaom
maintree and the evidence proves the production contract. - [~] Locally implemented, checkpoint open: production source exists, but current-tree verification is missing or a known audit issue invalidates the checkpoint.
- Remaining: the production behavior is absent, incomplete, or has not reached its required implementation boundary.
Current source reconciliation
Reconciled with the worktree on 2026-09-03.
- [~] The bounded container reader, still-image path, sequence parser, AV1 decoder, color pipeline, presentation pipeline, and broad AV1 test suite exist locally.
- The inter-frame decoder has verified checkpoints through inter deblocking decisions and reference/mode deltas.
- [~] Loop filtering, CDEF, super-resolution, restoration, film grain, layered presentation, alpha composition, and color conversion exist locally. Shared-source cleanup changed the current tree, so final production-path verification is open.
- [~] AV1 writer primitives, forward transforms, symbol encoding, and tile-writing source are connected to the public encoder for bounded still-image AVIF color and optional auxiliary alpha output.
- [~] The public AV1 encoder supports explicit single-image requests. Lossless output, bounded sequences, grids, orientation handling, and default format registration remain open.
- Patented codec production code, registrations, tests, benchmarks, fixtures, reference outputs, and notices were manually deleted and committed by
78a74d448. - Remaining task-created HM, HEVC, libheif, GPAC, Nokia, FFmpeg, Pillow HEIF, libavif-build, and libjpeg-build directories were traced to their creation commands in the recovered Codex session history and deleted on 2026-08-31. The user-provided repositories and all libaom-only source, build, and reference data were left untouched.
- The PNG metadata-suppression fix and three HEIF/AV1 diagnostic-save call-site corrections passed the exact 34 net11.0 ARM CI cases and were committed with the single-reference checkpoint as
54bb6cbe59bd113058854a3ee31448cf61f462ca. They are infrastructure evidence, not decoder or encoder completion evidence. - The complete decoder and encoder release matrix is not complete.
Immediate execution queue
Work must proceed in this order. Do not skip to a later item while an earlier checkpoint is open.
1. Finish and verify the AV1-only cleanup
- Remove production types, registrations, constants, parser branches, properties, tests, benchmarks, fixtures, reference outputs, notices, and documentation for removed codec work.
- Remove downloaded non-libaom reference source, tools, generated outputs, and local installations.
- Retain the official current-main libaom checkout and libaom-only build artifacts required for AV1 verification.
- Retain user-supplied AV1 fixtures and their recorded expected outputs.
- Audit production source, tests, benchmarks, assets, project files, notices, and documentation for stale removed-code references.
- The cleanup and cICP tree built in Release for net10.0 and net11.0 with restore disabled, build servers disabled, and one MSBuild node.
- The exact 34 net11.0 ARM CI failures pass after the cICP correction, and the subsequent single-reference checkpoint set passes on net10.0 and net11.0.
- Roslynk, scoped StyleCop, whitespace, and
git diff --checkaccepted the cleanup and cICP checkpoint. - The cleanup and cICP evidence was recorded and committed with the single-reference checkpoint.
Historical cleanup evidence from 2026-08-30, retained with its limitation:
- Release source builds passed for net10.0 and net11.0 with zero warnings and zero errors. Both builds used
--no-restore,--disable-build-servers, and one MSBuild node. - The focused net10.0 HEIF decoder, encoder, metadata, sequence-parser, and AV1 reconstruction set passed 221 of 221 tests with zero failures and zero skips. It did not execute the net11.0 diagnostic-save path that later failed in CI.
- The Roslyn compiler and configured StyleCop analyzers accepted the changed production source. Roslynk's
open_solutionentry point was attempted separately but failed before returning a solution handle, so no Roslynk result is claimed. - The tracked-source text and filename audit found no removed-code references outside the unchanged repository and shared-infrastructure
.gitattributespatterns. A later history reconstruction found ignored task-created reference directories that this audit missed; those directories were deleted on 2026-08-31. git diff --checkpassed and neither.gitattributesfile changed.
Current cICP failure correction evidence from 2026-08-31:
- The failure was not decoded HEIF metadata.
PngEncoderCore.WriteCicpChunkignoredPngChunkFilter.ExcludeAll, so diagnostic PNG saves attempted to write a non-identity source matrix that PNG cannot represent. PngEncoderCorenow honors the existingSkipMetadatacontract for cICP, and the three affected HEIF/AV1 diagnostic saves explicitly usePngEncoder { SkipMetadata = true }. Actual comparisons and decoded-image metadata assertions remain unchanged.- The direct embedded-ICC case and every row of the 12-case profile matrix passed: 13 of 13 net11.0 Release cases.
- The exact 34 cases reported by CI passed: 34 of 34 net11.0 Release cases, with zero failures and zero skips.
- Roslynk reported zero compiler errors after the fix, and
git diff --checkpassed. - The remaining four branch-introduced
GC.AllocateUninitializedArraycalls are removed from AV1 configuration, pixel-information, XMP, and Exif ownership boundaries. Each retained value still receives exactly one array and one copy because its source span belongs to pooled storage; no second materialization was introduced. The exact net11 Release rebuild remains at 1,005 warnings and zero errors, 166 focused configuration and metadata cases pass, and all 9,181 HEIF tests pass through direct VSTest.
Recovered task-history evidence from 2026-08-31:
- The primary session beginning on 2026-08-24 was reopened from task ID
01a03239-831b-7831-84e7-7f6947279ccb: 96,777 records, 295 turn contexts, 211 compactions, 190 user messages, 1,920 assistant messages, and 13,671 tool calls. - The continuation beginning on 2026-08-27 was reopened from task ID
01a04314-f1c6-7133-b1bc-5c74a94dd714: 61,129 records at the audit point, 166 turn contexts, 96 compactions, 223 user messages, 1,113 assistant messages, and 8,942 tool calls. - The restored first session records the user selecting official AOM/libaom as the AV1 source after the ImageSharp discussion was inspected. It does not authorize another codec implementation as an AV1 source and does not authorize importing a patented codec.
- The restored tool calls identify the exact creation commands for the non-libaom source, tool, and output directories removed on 2026-08-31. No directory was selected for deletion from its name alone.
- The recovered Git sequence establishes that
78a74d448removed the patented codec implementation and92fa7a8camerged the later upstream ImageSharp changes. The current branch and worktree, not an older summary, remain authoritative.
2. Correct the single-reference inter-frame checkpoint
The checkpoint is complete through c4b4e4e0386328dea574a884b6fa36c360ad5a9b. It replaces frame-sized palette maps with fixed decoder-session scratch, reconstructs each superblock before reusing that scratch, and passes the ownership, documentation, full AV1 test, and Release source-build gates on both target frameworks.
- Reconcile interpolation-filter syntax in
Av1TileReaderwith current libaommain.- Current libaom
av1_is_interp_neededcallsis_nontrans_global_motion, whose loop rejects onlyTRANSLATION. Identity GLOBALMV therefore omits switchable-filter symbols. - Current
Av1TileReaderuses the same non-Translation classification. The existing Identity test leaves sentinel filter symbols unread, while the Translation test consumes them. - No production change is required. The focused test describes only the syntax behavior it proves.
- Current libaom
- Reconcile both spatial single-reference extension loops in
Av1ReferenceMotionVectorswith current libaommain.- Current libaom
setup_ref_mv_liststops both loops atMAX_MV_REF_CANDIDATES, which is two.MAX_REF_MV_STACK_SIZE, which is eight, is the stack capacity used by the earlier direct and temporal candidate collection; it is not the stop condition for these two extension loops. - Current
Av1ReferenceMotionVectorsuses the same two-entry stop condition and retains an eight-entry stack for earlier candidates and DRL selection. - No production change is required. This remains spatial single-reference extension, not temporal extension.
- Current libaom
- Establish and enforce the contiguous frame-plane invariant used by
Av1FrameBufferand inter reconstruction.- One ImageSharp allocator owner now contains the aligned Y, U, and V storage, matching libaom's frame-buffer ownership while non-owning
Buffer2Dviews preserve ImageSharp's row API. Coded dimensions are aligned to eight samples, the luma stride is aligned to 32 samples, and chroma strides and heights are derived from that luma layout exactly once. A 4K eight-bit 4:2:0 frame owner occupies about 17.3 MiB. - The single owner removes the previous three-rent constructor and its allocation-cleanup
try/catch.Av1FrameBufferrejects external geometry whose complete aligned frame reaches the contiguousint.MaxValueboundary before allocation, making every directDangerousGetSingleSpancall an enforced owner invariant. ConstructorRequestsContiguousPaddedPlanesproves that a frame larger than the allocator's group capacity remains one group.ConstructorUsesOneFrameOwnerForAllPaddedPlanesproves exact one-rent Y/U/V ownership and exactly-once return.ConstructorRejectsPaddedPlaneThatCannotBeContiguousproves that an unrepresentable frame is rejected before allocation, and the high-bit-depth stride regression proves the 608-sample libaom layout for a three-pixel coded row.- The complete HEIF/AV1 namespace passes 8,808 of 8,808 direct net11 VSTest cases in Release after the physical layout change. The production path performs no plane copy and no per-block, per-row, or per-scanline allocation.
- One ImageSharp allocator owner now contains the aligned Y, U, and V storage, matching libaom's frame-buffer ownership while non-owning
- Prove the real
Av1BlockDecoder.DecodeBlockinter-reconstruction branch.- Decode the progressive dependent-frame fixture through the complete public production path.
- Compare the final frame's native Y, Cb, and Cr planes exactly with current-main libaom output.
- Compare the final presented image through the established ImageSharp reference-image comparison API.
- Do not substitute an internal helper test, fake tile reader, non-zero assertion, custom pixel loop, or tolerant comparison.
- Prove motion-field ownership and lifetime after the current reconstruction-timing change.
- Track initialization, retained-slot aliases, failure unwinding, presentation ownership, decoder-result ownership, and final disposal.
- Every allocator-owned object must be returned exactly once.
- Correct stale documentation for the current worktree.
- Av1InterFrameModeInfoTests must describe the behavior it actually proves.
- Do not claim production reconstruction, constrained allocation, ownership, or reference-stack coverage unless the test executes that contract.
Checkpoint gate:
- Default Identity-omission and Translation-consumption GLOBALMV syntax cases pass in the focused current-tree run.
- Two-entry spatial single-reference extension passes; current-main source inspection confirms the separate eight-entry overall stack capacity and DRL access.
- The exact dependent-frame native-plane comparison passes.
- The established exact presentation comparison passes.
- Normal, AVX-512-disabled, AVX-disabled, and scalar FeatureTestRunner configurations pass where supported.
- Constrained allocation preserves the enforced single-group plane invariant without copying or per-block allocation.
- Motion-field allocation tracking is balanced across success and failure on net10.0 and net11.0.
- Release source builds pass for net10.0 and net11.0 with zero warnings and zero errors.
- The complete AV1 namespace passes 8,732 of 8,732 tests on net10.0 and net11.0 with zero failures or skips.
- Roslynk reports zero compiler errors; scoped analyzer inspection reports no diagnostics introduced by the current changes;
git diff --checkpasses. - The completed checkpoint was committed as
54bb6cbe59bd113058854a3ee31448cf61f462cawith author and committerJames Jackson-South <james_south@hotmail.com>. - The palette-memory follow-up was committed as
c4b4e4e0386328dea574a884b6fa36c360ad5a9bwith author and committerJames Jackson-South <james_south@hotmail.com>.
Verified single-reference checkpoint evidence on 2026-08-31:
- The current-main
aomdecwas rebuilt directly fromD:\GitHub\AOMediaCodec\aomand identified itself as3.15.0-13-g441c439b99. - Decoding the 72-byte progressive payload with
--all-layers, one thread, and row multithreading disabled produced 2,178 YUV444 color samples. All samples in both layers match the first three planes of the stored YUV444-alpha reference exactly. DecodeProgressiveSingleMatchesReferenceexecutes the production decoder through FeatureTestRunner and compares the complete presentedRgba32image withCompareToReferenceOutput(ImageComparer.Exact, provider). The redundant manual alpha loop was removed.DecodeProgressiveSingleWithConstrainedAllocatorexecutes the same production reconstruction with a 1,024-byte allocator group capacity and verifies that every allocation is returned exactly once.MotionFieldsFollowAliasesAndPresentationOwnership,MotionFieldAllocationFailureUnwindsTileReaderOwnership,DecodeProgressiveSingleTracksMotionFieldOwnership, and the reference-store replacement, reset, and transfer tests cover initialization, aliases, presentation ownership, decoder-result ownership, failure unwinding, repeated disposal, and exactly-once final returns in the current worktree.- The current worktree passes the four-case palette set, seven-case ownership set, and 29-case syntax, plane, and production reconstruction set on both target frameworks. The complete AV1 namespace passes 8,732 of 8,732 tests on net10.0 and net11.0 with zero failures or skips.
- Release source builds passed for net10.0 and net11.0 with zero warnings and zero errors.
- Roslynk reported zero compiler errors. The scoped changed-file analyzer inspection reported no StyleCop diagnostics attributable to this checkpoint; its only remaining match is the pre-existing xUnit cancellation warning in an unrelated
HeifDecoderTestsmethod. git diff --checkpassed, and neither.gitattributesfile changed.
Exact verification commands, run directly in the foreground from D:\GitHub\ynse01\ImageSharp:
$env:MSBUILDUSESERVER = '0'
$env:DOTNET_CLI_USE_MSBUILD_SERVER = '0'
$env:DOTNET_CLI_HOME = 'D:\GitHub\ynse01\ImageSharp\.dotnet'
$env:DOTNET_SKIP_FIRST_TIME_EXPERIENCE = '1'
$env:DOTNET_CLI_TELEMETRY_OPTOUT = '1'
$env:DOTNET_DbgEnableMiniDump = '0'
$env:COMPlus_DbgEnableMiniDump = '0'
$env:DOTNET_EnableCrashReport = '0'
$env:COMPlus_EnableCrashReport = '0'
$heifCheckpointFilter = 'FullyQualifiedName~Av1InterFrameModeInfoTests.ReadInterFrameModeInfoReadsInterpolationFilters|FullyQualifiedName~Av1InterFrameModeInfoTests.IdentityGlobalMotionOmitsInterpolationFilters|FullyQualifiedName~Av1ReferenceMotionVectorsTests.BuildReversesOppositeDirectionExtensionCandidate|FullyQualifiedName~Av1FrameBufferTests|FullyQualifiedName~Av1ReferenceFrameStoreTests.MotionFieldsFollowAliasesAndPresentationOwnership|FullyQualifiedName~Av1ReferenceFrameStoreTests.MotionFieldAllocationFailureUnwindsTileReaderOwnership|FullyQualifiedName~Av1ReferenceFrameStoreTests.PartialReplacementPreservesSharedOwner|FullyQualifiedName~Av1ReferenceFrameStoreTests.FinalReplacementReleasesDisplacedOwner|FullyQualifiedName~Av1ReferenceFrameStoreTests.ResetReleasesUniqueOwnersAndClearsSlots|FullyQualifiedName~Av1ReferenceFrameStoreTests.TakeOutputTransfersPlanesAndReleasesOtherReferences|FullyQualifiedName~Av1ReconstructionConformanceTests.DecodeProgressiveSingleMatchesReference|FullyQualifiedName~Av1ReconstructionConformanceTests.DecodeProgressiveSingleWithConstrainedAllocator|FullyQualifiedName~Av1ReconstructionConformanceTests.DecodeProgressiveSingleTracksMotionFieldOwnership'
dotnet build src\ImageSharp\ImageSharp.csproj -c Release -f net10.0 --no-restore --disable-build-servers -m:1 --no-incremental --nologo --verbosity:minimal
dotnet build src\ImageSharp\ImageSharp.csproj -c Release -f net11.0 --no-restore --disable-build-servers -m:1 --no-incremental --nologo --verbosity:minimal
dotnet test tests\ImageSharp.Tests\ImageSharp.Tests.csproj -c Release -f net10.0 --no-restore --disable-build-servers -m:1 --filter $heifCheckpointFilter --logger 'console;verbosity=minimal'
dotnet test tests\ImageSharp.Tests\ImageSharp.Tests.csproj -c Release -f net11.0 --no-restore --disable-build-servers -m:1 --filter $heifCheckpointFilter --logger 'console;verbosity=minimal'
$aomVcVars = 'C:\Program Files (x86)\Microsoft Visual Studio\18\BuildTools\VC\Auxiliary\Build\vcvars64.bat'
$aomCmake = 'C:\Program Files (x86)\Microsoft Visual Studio\18\BuildTools\Common7\IDE\CommonExtensions\Microsoft\CMake\CMake\bin\cmake.exe'
$aomEnvironment = & cmd.exe /d /s /c "`"$aomVcVars`" >nul && set"
foreach ($aomEntry in $aomEnvironment)
{
$aomParts = $aomEntry -split '=', 2
if ($aomParts.Length -eq 2)
{
[Environment]::SetEnvironmentVariable($aomParts[0], $aomParts[1], 'Process')
}
}
& $aomCmake --build artifacts\reference\aom-generic --target aomdec --config Release --parallel 1
& 'artifacts\reference\aom-generic\aomdec.exe' --codec=av1 --rawvideo --all-layers --threads=1 --row-mt=0 --output='artifacts\reference\aom-generic\progressive-current-main-all-layers.yuv' 'tests\Images\Input\Heif\Av1\Conformance\libavif-progressive-draw-points-8b.bit'
3. Reverify downstream inter prediction in recorded order
The single-reference syntax, buffer, reconstruction, and ownership foundation is verified by 54bb6cbe59bd113058854a3ee31448cf61f462ca. Reverify the existing downstream implementations in this exact order, treating each as locally implemented but unverified until its current-main evidence is recorded.
- Compound reference selection, paired reference-MV derivation, and equal averaging.
- Inter-intra prediction.
- Distance-weighted compound prediction.
- Wedge compound prediction.
- Difference-weighted compound prediction.
- OBMC.
- Scaled-reference prediction.
- Local warped prediction.
- Non-translational global prediction.
- Inter deblocking decisions and reference/mode deltas.
Verified equal-average compound checkpoint evidence on 2026-08-31:
- Refreshed the clean official libaom
maincheckout and audited the observed revision441c439b9916474cac15d2822af47a9ad70674a8. Reference selection and compound mode syntax matchread_comp_reference_typeandread_ref_framesinav1/decoder/decodemv.c; contexts matchav1/common/pred_common.c; paired reference-MV construction and eight-entry extension matchprocess_compound_ref_mv_candidateandsetup_ref_mv_listinav1/common/mvref_common.c. - Audited equal-average reconstruction against
av1/common/convolve.candav1/common/convolve.h. Corrected the unscaled 10/12-bit translational path so both references retain libaom's no-round compound intermediates until the sole final average and clipping step, including the larger first-round shift required for 12-bit horizontal intermediates. - Added descending Vector512, Vector256, Vector128, and scalar high-bit-depth traversal to the existing semantic compound-prediction operator families. No per-block, per-row, or per-scanline allocation or copy was added.
- Added FeatureTestRunner coverage for 10/12-bit copy, horizontal, vertical, and separable subpixel prediction at widths 9, 17, 33, and 65, with an independent no-round bilinear oracle, row-padding sentinels, and explicit scalar comparison.
- Added a complete
Av1BlockDecoder.DecodeBlock10/12-bit half-sample regression whose expected result comes from the scalar no-round pipeline. The selected vector differs by one sample from the obsolete round-each-reference behavior, so the test proves the production branch selection. - Refreshed the official libaom
mainremote immediately before verification and decoded the fixture's 5,465-byte AV1mdatpayload with currentaomdec, one thread and row threading disabled. All 19 frames decoded; the final 19,200 YUV444 samples have SHA-256E79D2F49C260B1AC9B1B9BBBB2D611126AFD3B241DA389EB9E7BD4EA0ED42080and match the retained native reference with zero differing samples. - The real 19-frame production sequence requires decoded equal-average compound blocks, compares the final native Y, U, and V planes exactly, compares final RGBA presentation through ImageSharp's established reference-output API, and repeats the complete decode with a 1,024-byte constrained tracked allocator and exactly-once return checks.
- The focused Release checkpoint set passes 31/31 on net10.0 and 31/31 on net11.0, with zero
failures or skips. Scoped analyzer and whitespace verification pass for every changed C# file,
Roslynk reports zero compiler errors and no diagnostics in the changed files, and
git diff --checkpasses..gitattributesis unchanged. - The completed checkpoint was committed as
4075a0844836e863a93cb2e2f3ca42d202c7df1bwith author and committerJames Jackson-South <james_south@hotmail.com>.
Verified inter-intra checkpoint evidence on 2026-08-31:
- Audited syntax against current libaom
av1/decoder/decodemv.candav1/common/blockd.h. ImageSharp applies the same sequence enable, skip-mode, block-size, and single-reference gates, reads the same four-mode CDF, and reads wedge syntax only within libaom's wedge-supportedBLOCK_8X8throughBLOCK_32X32range. - Audited reconstruction against
ii_weights1d,ii_size_scales,build_smooth_interintra_mask, andcombine_interintrain currentav1/common/reconinter.c. The ImageSharp weights, plane-size scaling, smooth-mask direction, complemented destination orientation, wedge sign, subsampling, and final 6-bit blend match. No production change was required. - The mask tests cover all four inter-intra modes, complemented orientation, row-padding
sentinels, and the 32-wide curve. FeatureTestRunner covers byte and high-bit-depth selectable
blending under SIMD and scalar dispatch, and complete
Av1BlockDecoder.DecodeBlocktests execute smooth inter-intra reconstruction at 8, 10, and 12 bits. - Extracted the fixture's 5,327-byte AV1
mdatpayload and decoded it with the refreshed current libaomaomdec, using one thread with row threading disabled. All 19 frames decoded. The final 19,200 YUV444 samples have SHA-256E8B776C2751DC30CA838931A4B74535FC6E681179568A1278747A38CFF2E5BFAand match the retained native reference with zero differing samples. - The real production sequence requires both smooth and wedge inter-intra blocks, compares the final native Y, Cb, and Cr planes exactly, and compares final RGBA presentation through ImageSharp's established reference-output API. Its constrained 1,024-byte tracked-allocator run proves motion-field allocation and exactly one return for every allocation.
- The focused Release checkpoint set passes 50/50 on net10.0 and 50/50 on net11.0, with zero
failures or skips. Scoped analyzer and whitespace verification pass for both changed C# files.
Roslynk reports zero compiler errors and no diagnostics in the changed files,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
18b1c881271a3494489ca6f410ab140544902e2dwith author and committerJames Jackson-South <james_south@hotmail.com>.
Verified distance-weighted compound checkpoint evidence on 2026-08-31:
- Audited reference-distance quantization against
quant_dist_weightandquant_dist_lookup_tablein current libaomav1/common/common_data.h, and audited order-hint distance selection and forward/backward reference assignment againstav1_dist_wtd_comp_weight_assigninav1/common/reconinter.c. - Audited reconstruction against current libaom
av1/common/convolve.c. Corrected the production 10/12-bit subpixel path, which incorrectly finalized its two no-round compound intermediates with an equal average instead of the signaled distance weights. The fixed path applies libaom's 4-bit weighted shift before bias removal, final rounding, and clipping. - Added descending Vector512, Vector256, Vector128, and scalar traversal to the existing semantic distance-weighted intermediate predictor family. Unsigned widening preserves the biased 12-bit intermediate range. No per-block, per-row, or per-scanline allocation or copy was added.
- Added FeatureTestRunner coverage for every current-libaom distance-weight class in both reference orders, and for 10/12-bit copy, horizontal, vertical, and separable subpixel prediction at widths 9, 17, 33, and 65, with an independent no-round oracle and row-padding sentinels.
- Added a complete
Av1BlockDecoder.DecodeBlock10/12-bit half-sample regression that selects the 13:3 distance weights through real order hints. Its first reconstructed sample differs from the old equal-average result, so the test proves the corrected production branch is executed. - Extracted the fixture's 5,372-byte AV1
mdatpayload and decoded it with the refreshed current libaomaomdec, using one thread with row threading disabled. All 19 frames decoded. The final 19,200 YUV444 samples have SHA-256E8CAA650F1571C5B9CACAF8C06E1DDF5F5D2ED35F65F1C34377076C573425899and match the retained native reference with zero differing samples. - The real 19-frame production sequence requires decoded distance-weighted compound blocks, compares the final native Y, Cb, and Cr planes exactly, compares final RGBA presentation through ImageSharp's established reference-output API, and repeats the complete decode with a 1,024-byte constrained tracked allocator and exactly-once return checks.
- The focused Release checkpoint set passes 44/44 on net10.0 and 44/44 on net11.0, with zero failures
or skips. Scoped analyzer and whitespace verification pass for every changed C# file. Roslynk reports
zero compiler errors and no diagnostics in the changed files,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
7e2de7a2c25852acc374b17936a1a644464f77f3with author and committerJames Jackson-South <james_south@hotmail.com>.
Verified wedge compound checkpoint evidence on 2026-08-31:
- Audited mask generation against current libaom
tools/gen_wedge_masks_data.pyandav1/common/reconinter.c, including the master prototypes, direction transforms, block-size codebooks, sign flips, offsets, and luma/chroma mask sampling. ImageSharp's generated masks match those definitions; only stale “pinned” documentation required correction. - Audited reconstruction against current libaom
aom_dsp/blend_a64_mask.c. The high-bit-depth d16 path applies the Q6 mask to both no-round intermediates before bias removal, the sole final rounding step, and clipping. - Corrected the production high-bit-depth intermediate eligibility gate, which admitted only equal-average blocks and made the distance-weighted and wedge no-round finalizers unreachable. Average, distance-weighted, and wedge subpixel blocks now retain both intermediates until their signaled finalizer; difference-weighted blending remains excluded for its next ordered checkpoint.
- Added high-bit-depth traversal to the existing semantic mask-blend predictor and readonly operator family with descending Vector512, Vector256, Vector128, and scalar dispatch. Unsigned widening preserves the biased 12-bit intermediate range. No per-block, per-row, or per-scanline allocation or copy was added.
- Extended FeatureTestRunner coverage with an independent Q6 mask oracle across 10/12-bit copy,
horizontal, vertical, and separable subpixel prediction, widths 9, 17, 33, and 65, all mask weights
from 0 through 64, and row-padding sentinels. A complete
Av1BlockDecoder.DecodeBlockregression verifies the current-libaom 8x8 wedge mask and the production no-round branch. - Extracted the fixture's 5,374-byte AV1
mdatpayload and decoded it with refreshed current libaomaomdec, using one thread with row threading disabled. All 19 frames decoded. The final 19,200 YUV444 samples have SHA-256E8CAA650F1571C5B9CACAF8C06E1DDF5F5D2ED35F65F1C34377076C573425899and match the retained native reference with zero differing samples. - The real 19-frame production sequence requires both wedge-mask orientations, compares final native Y, Cb, and Cr planes exactly, compares final RGBA presentation through ImageSharp's established reference-output API, and repeats the complete decode with a 1,024-byte constrained tracked allocator and exactly-once return checks.
- The focused Release checkpoint set passes 35/35 on net10.0 and 35/35 on net11.0, with zero
failures or skips. Scoped analyzer and whitespace verification pass for every changed C# file.
Roslynk reports zero compiler errors,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
9883a24dc319e16b471f68be632d4f62f2c1cd5ewith author and committerJames Jackson-South <james_south@hotmail.com>.
Verified difference-weighted compound checkpoint evidence on 2026-08-31:
- Audited syntax against current libaom
av1/decoder/decodemv.c. ImageSharp applies the same masked-compound enable and block-size gates, selects difference-weighted compound directly when wedge is unavailable, and reads the same one-bit type-38 mask orientation. - Audited mask generation and reconstruction against current libaom
av1/common/reconinter.candaom_dsp/blend_a64_mask.c. The d16 path rounds the absolute intermediate difference by the convolution and bit-depth shift, scales it by 1/16, adds the type-38 base, clamps or inverts the mask, and then blends the original no-round intermediates before final rounding and clipping. Chroma reuses the luma-derived mask through rounded subsampling. - Corrected the production 10/12-bit subpixel eligibility gate, which previously rounded both references before difference-mask construction and blending. Difference-weighted blocks now use the existing semantic intermediate mask-builder and mask-blend predictor/operator families through the sole final rounding step. No new operator family, per-block allocation, or copy was introduced.
- Renamed the stale “pinned formula” test and extended FeatureTestRunner's independent oracle across current-libaom regular and d16 mask arithmetic, both mask orientations, 8/10/12-bit samples, widths that cross every Vector512, Vector256, Vector128, and scalar boundary, subpixel phases, and row-padding sentinels.
- Added a complete
Av1BlockDecoder.DecodeBlockregression for 10/12-bit half-sample prediction and both type-38 orientations. Its expected mask and reconstruction are calculated directly from the current-libaom equations, independently of the production mask builder and finalizer. - Extracted the fixture's 5,358-byte AV1
mdatpayload and decoded it with refreshed current libaomaomdec, using one thread with row threading disabled. All 19 frames decoded. The final 19,200 YUV444 samples have SHA-256E8CAA650F1571C5B9CACAF8C06E1DDF5F5D2ED35F65F1C34377076C573425899and match the retained native reference with zero differing samples. - The real 19-frame production sequence requires both difference-mask orientations, compares final native Y, Cb, and Cr planes exactly, compares final RGBA presentation through ImageSharp's established reference-output API, and repeats the complete decode with a 1,024-byte constrained tracked allocator and exactly-once return checks.
- The focused Release checkpoint set passes 37/37 on net10.0 and 37/37 on net11.0, with zero
failures or skips. Scoped analyzer and whitespace verification pass for every changed C# file.
Roslynk reports zero compiler errors,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
fb4c64474e1ced4067a42731384f3b5ad4212a2fwith author and committerJames Jackson-South <james_south@hotmail.com>.
Verified OBMC checkpoint evidence on 2026-08-31:
- Audited motion-mode syntax against current libaom
av1/decoder/decodemv.c,av1/common/blockd.h,av1/common/reconinter.c,av1/common/obmc.h, andav1/common/reconinter_template.inc. ImageSharp applies the same switchable-mode, skip, single-reference, inter-intra, minimum-size, overlappable-neighbor, fixed-global-motion, scaled reference, and projection-sample gates and reads the matching binary or three-way CDF. - Audited above and left neighbor traversal, 4x4 pairing, neighbor caps, chroma suppression, prediction rectangles, interpolation filters, first-reference selection, mask tables, and blend order against current libaom. The existing semantic mask-blend predictor remains the correct SIMD-first traversal; no OBMC-specific operator family, allocation, or copy was introduced.
- Corrected the unscaled neighbor far-edge UMV clamp. After converting libaom's neighbor-relative motion-vector limits to an absolute source coordinate, the prediction extent cancels from the right and bottom limits; the previous code counted it twice.
- Extracted the fixture's 5,387-byte AV1
mdatpayload at AVIF offset 1,065 and decoded it with refreshed current libaomaomdec, using one thread with row threading disabled. All 19 frames decoded. The final 19,200 YUV444 samples have SHA-256E8CAA650F1571C5B9CACAF8C06E1DDF5F5D2ED35F65F1C34377076C573425899and match the retained native reference with zero differing samples. - The production sequence asserts decoded OBMC mode state, compares final native Y, Cb, and Cr
planes exactly, compares final RGBA presentation through ImageSharp's established reference-output
API under normal and scalar FeatureTestRunner dispatch, and repeats reconstruction with a 1,024-byte
constrained tracked allocator. Direct
DecodeBlocktests cover above-then-left blending at 8/10/12-bit and 4:2:0 and 4:2:2 chroma geometry. - Renamed the stale pinned-reference test and its established reference-output PNG together. The
PNG SHA-256 remains
D2CB388C9092EF17C4F0382C0150DD30D6F9D0EE247FF45AB5D7D4D312CEB23C; only its contract-derived filename changed. - The focused Release checkpoint set passes 18/18 on net10.0 and 18/18 on net11.0, with zero
failures or skips. Scoped analyzer and whitespace verification pass for every changed C# file.
Roslynk reports zero compiler errors,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
7e7e3cbe6438d63926b31d966795d2652e221939with author and committerJames Jackson-South <james_south@hotmail.com>.
Verified scaled-reference checkpoint evidence on 2026-08-31:
- Audited reference-size validation and variable-scale coordinates, filters, edge extension, convolution
rounding, and compound intermediates against current libaom
av1/common/scale.c,av1/decoder/decodeframe.c, andav1/common/convolve.c. The frame boundary accepts the same half-to-sixteen-times dimension range and requires at least one compatible selected reference. - Corrected the production scaled-compound branch. It previously rounded each scaled reference into
native pixels before blending; current libaom retains both
CONV_BUF_TYPEvalues withCOMPOUND_ROUND1_BITSequal to seven and performs one final rounding after the selected compound blend. - Kept native-pixel and compound output in the existing
Av1ScaledInterPredictortraversal with semanticNativeOperatorandCompoundOperatoroutput contracts. The closed generic traversal shares variable-phase arithmetic across byte and ushort sources, dispatches Vector512, Vector256, Vector128, then scalar, and adds no per-block allocation or copy. - Added independent FeatureTestRunner oracles for native and no-round compound output across 8, 10,
and 12 bits, variable phases, all interpolation families, reduced kernels, vector tails, and destination
padding. A complete
Av1BlockDecoder.DecodeBlock()regression covers scaled compound prediction across all, AVX-512-disabled, AVX-disabled, and scalar configurations and proves the vector differs from an incorrectly early-rounded blend. - Decoded the 2,195-byte layered payload with refreshed current libaom
aomdec, using one thread, row threading disabled, all layers selected, and raw 8-bit output. The 40x40 YUV444 base and 80x80 YUV444 dependent frames total 24,000 samples with SHA-256DD219E41B52C6C9343A92CD0A2D451DF57B73B25F10124811675B4CB2F8D666F; both match their retained native references with zero differing samples. - The production tests compare both native frames exactly, compare selected-layer and final RGBA presentation through ImageSharp's established reference-output API, and repeat both paths with a 1,024-byte constrained tracked allocator whose allocations have balanced exactly-once returns.
- Renamed the two stale pinned-reference tests and their contract-derived PNGs together. Their Git blob
identifiers remain unchanged, and their SHA-256 values remain
DC4C6DBE6BD92C5FCE1E3E23700AFA603EF04ED02EDD336213EBBA1E3BD84BA0and678C5E5D4650EA6F0C590302E7DB9E3C6608851BC577453DA4A6837BDB4D3AF3. - The focused Release checkpoint set passes 10/10 on net10.0 and 10/10 on net11.0, with zero failures
or skips. Scoped analyzer and whitespace verification pass for every changed C# file. Roslynk reports
zero compiler errors,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
658a9cd1b6e22806decbae923da8800bca03a09ewith author and committerJames Jackson-South <james_south@hotmail.com>.
Verified local warped-prediction checkpoint evidence on 2026-08-31:
- Refreshed the clean official libaom
maincheckout and audited the observed revision441c439b9916474cac15d2822af47a9ad70674a8. Motion-mode eligibility and CDF selection matchread_motion_modeinav1/decoder/decodemv.c; above, left, top-left, and top-right spatial projection samples and threshold selection matchfindSamplesandselectSamplesinav1/common/mvref_common.c; affine fitting, shear reduction, phase derivation, filters, rounding, clipping, and invalid-model fallback matchav1/common/warped_motion.candav1/common/reconinter.c. - Mechanically compared all 1,544 ImageSharp and independent-test warped-filter coefficients against
current libaom's
av1_warped_filter; both comparisons have zero differences. The separate scalar test transcription covers 8-, 10-, and 12-bit luma and subsampled-chroma coordinates, tail widths, destination stride preservation, libaom's 12-bit round adjustment, and AVX-512, AVX, 128-bit, and scalar dispatch throughFeatureTestRunner. - Extracted the fixture's exact 2,310-byte AV1
mdatpayload at AVIF offset 997. Its SHA-256 is644D04FE1D1A32BB7A3856AD7EB49CF1EFDE0AC845E55BEAC4170F72353F2391. Current official libaom decoded both 256x256 YUV444 frames with one thread, row threading disabled, and all layers enabled. The complete Y4M SHA-256 is8FDC5D46014F5E5A7455A83643AB6F0DA66FC5A984E72A43F8C75BAD8271C299; the final frame's 196,608 native samples have SHA-25647B2AB39BF3B9DA15C1EC59840F964DFDF227760947F6E1295FB38A84555F75Cand match the retained native reference with zero differences. - The real two-frame fixture exercises
Av1BlockDecoder.DecodeBlock(), requires decodedWARPED_CAUSALstate and the expected multi-sample affine model, compares final native Y, U, and V planes exactly, compares the retained final presentation through ImageSharp's established reference-output API, and passes through intrinsic and scalar dispatch. The 1,024-byte constrained tracked-allocator path passes with motion-field allocations present and balanced exactly-once returns. - Renamed the stale pinned-reference test and its contract-derived PNG together without changing the PNG
bytes. Its SHA-256 remains
4490D62FB6679378E92CACA48427359091AD2106BE49FC1A3848F78BE03BEEB1. - The focused Release checkpoint set passes 4/4 on net10.0 and 4/4 on net11.0, with zero failures or
skips. Scoped analyzer verification passes for both changed C# files. Roslynk reports zero compiler errors,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
27a522424fe7aaea25078e705d71a501da110727with author and committerJames Jackson-South <james_south@hotmail.com>.
Verified non-translational global-prediction checkpoint evidence on 2026-08-31:
- Audited global-motion syntax, coefficient decoding, previous-reference recentering, shear validation,
motion-vector projection, and warped-prediction eligibility against current official libaom
mainat the observed revision441c439b9916474cac15d2822af47a9ad70674a8. The implementation matchesread_global_motion_params,read_global_motion_model,gm_get_motion_vector,is_global_mv_block, and the WARP_PRED selection inav1/common/reconinter.c. - Corrected high-bit-depth compound warped/global prediction to retain both references in libaom's
unsigned no-round compound domain. Current
get_conv_params_no_round,av1_warp_plane, andav1_highbd_warp_affine_crequire the 12-bit first-round adjustment while retaining a seven-bit second round; native clipping now occurs only after the compound blend. - The independent scalar libaom transcription validates native and no-round compound output for byte,
8-bit, 10-bit, and 12-bit sources, including tail widths and destination-stride preservation. All cases pass
through AVX-512, AVX, 128-bit, and scalar dispatch with
FeatureTestRunner. DirectAv1BlockDecoder.DecodeBlock()coverage validatesGLOBAL_GLOBALMVcompound reconstruction at all supported bit depths. - Extracted the fixture's exact 38,475-byte AV1
mdatpayload at AVIF offset 997. Its SHA-256 is6AC7EC9984B1FF5C00403D7E3858441E9CEE75128F7414101D06DEEE59A351D0. Current official libaom decoded both 256x256 YUV444 frames with one thread, row threading disabled, and all layers enabled. The complete Y4M SHA-256 is84754DE0B9FABC4F3F8F344C848183EC17B625BFD87E4519C3D8AD7DEFD20F2C; the final frame's 196,608 native samples have SHA-256FEC89E2DE7496980389806B194425042F3800C7BAA817249D1A51D44A2B37A8Eand match the retained native reference with zero differences. - The real two-frame fixture exercises the production decoder, requires decoded non-translational global motion, compares final native Y, U, and V planes exactly, compares the retained presentation through ImageSharp's established reference-output API, and passes the constrained tracked-allocator path.
- Renamed the stale pinned-reference test and its contract-derived PNG together without changing the PNG
bytes. Its SHA-256 remains
F7D27ABF79450DFA311F72106FD1DA80997EABC0937F2F5578EF627119FF83B0, and Git attributes select the LFS filter and diff driver. - The focused Release checkpoint set passes 11/11 on net10.0 and 11/11 on net11.0, with zero failures or
skips. Scoped analyzer verification passes for all six changed C# files. Roslynk reports zero compiler
errors,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
25295683d39a2336e9b98484c9fd54f33107ea66with author and committerJames Jackson-South <james_south@hotmail.com>.
Verified inter-deblocking checkpoint evidence on 2026-08-31:
- Audited frame-level loop-filter syntax and primary-reference inheritance against
setup_loopfilterin currentav1/decoder/decodeframe.c; per-superblock delta-LF parsing and prediction againstread_delta_q_paramsinav1/decoder/decodemv.c; and default reference/mode deltas againstav1/common/entropymode.cat observed current-main revision441c439b9916474cac15d2822af47a9ad70674a8. - Audited filter-level derivation, segmentation adjustment, reference scaling, global/non-global
mode classes, skipped-transform prediction-unit decisions, transform-edge selection, kernel length,
sharpness limits, and vertical-then-horizontal traversal against
get_filter_level,set_lpf_parameters,av1_filter_block_plane_vert,av1_filter_block_plane_horz, andav1_thread_loop_filter_rows. No production arithmetic change was required. - Added direct production
Av1LoopFilterDecoder.DecodeFrame()coverage using adjacent skipped 16x8 inter blocks split into 8x8 transforms. An independent scalar oracle proves that internal transform edges remain untouched and the prediction-unit edge uses current-libaom levels 17 for LAST/GLOBALMV, 21 for LAST/NEWMV, and 22 for GOLDEN/GLOBALMV. ExistingFeatureTestRunnercoverage continues to verify every filter width at 8, 10, and 12 bits under intrinsic and scalar dispatch. - Current official libaom decoded the retained 20,750-byte 8-bit, 37,169-byte 10-bit, and
23,769-byte 12-bit elementary streams with one thread, row threading disabled, raw output, and their
native output depths. The generated native files match the retained references byte for byte. Their
output SHA-256 values are
8DDE2EEC742C39F0579C29AE84CBA0FE01522A9008ADCB2CFFCCEC0295D18141,9A59DD92A0C579F942ACCA8281EBD0465DC848BE200A4D2FF57EAFF589445F6C, andEF712BE32AF7CF0A95C5C41BDCC51AFC05A4AB7C047383F5F65EDAD2BB986712. - Reused the already current-main scaled-reference sequence as the real inter checkpoint. It requires an inter frame with reference/mode-delta processing enabled, nonzero chroma filter levels, intra, inter, and skipped-inter blocks; compares both decoded native frames exactly; compares final presentation through ImageSharp's established reference-output API; and passes constrained tracked allocation with balanced returns.
- Removed an obsolete SVT-AV1 design link from mode-map documentation. Current official libaom remains the sole external codec implementation source.
- The focused Release checkpoint set passes 6/6 on net10.0 and 6/6 on net11.0, with zero failures or skips.
- Scoped analyzer verification passes for all four changed C# files. Roslynk reports zero compiler
errors,
git diff --checkpasses, and.gitattributesis unchanged. - The completed checkpoint was committed as
fcb502e4960cc7b8efb06b6f060e2c73a913a2bfwith author and committerJames Jackson-South <james_south@hotmail.com>.
For every item:
- Trace syntax and arithmetic to the current libaom
maintree. - Execute the real production decoder path.
- Compare native planes exactly.
- Compare presentation through the established reference-image API.
- Run constrained allocator and exactly-once ownership coverage.
- Run FeatureTestRunner for SIMD and scalar dispatch when the implementation has SIMD.
- Record focused Release evidence before marking the item verified.
4. Close AV1 decoder coverage
Previously verified algorithm checkpoints remain valuable evidence, but the final decoder gate requires a fresh current-tree run after the inter and cleanup corrections.
- Bounded OBU framing, sequence headers, frame headers, tile groups, alignment, and trailing-bit parsing have been re-audited and verified against current libaom
main. - Partition traversal, mode information, segmentation, delta quantization, transform-size selection, coefficient decoding, inverse quantization, and inverse transforms have been re-audited and verified against current libaom
main. - Intra prediction covers directional, DC, smooth, Paeth, chroma-from-luma, filter-intra, and palette families with the established operator architecture.
- Intra-block copy has exact native reconstruction and feature-isolated SIMD evidence.
- Lossless inverse transform, loop filtering, CDEF, super-resolution, restoration, and film grain have focused checkpoint evidence.
- Retained references, CDF snapshots, segmentation maps, global motion, temporal motion fields, and dependent-frame lifecycle have been re-audited and verified against current libaom
main. - The 12-case all-intra profile matrix covers every valid 8, 10, and 12-bit monochrome, 4:2:0, 4:2:2, and 4:4:4 combination. Dependent-frame coverage is recorded separately above.
- The exact current-tree native-plane matrix passes through the production decoder on net10.0 and net11.0. The normal-dispatch and FeatureTestRunner fallback methods pass 2 of 2 focused tests on each target.
- The exact current-tree presentation matrix passes 12 of 12 cases through ImageSharp's established reference-image API on net10.0 and net11.0.
- Verify malformed/truncated data, frame IDs, reference slots, tile bounds, allocation limits, cancellation, and failure unwinding.
- Verify still items and bounded sequences from file, memory, non-seekable, and short-read streams.
- Verify ICC, CICP, alpha, grids, pixel aspect ratio, clean aperture, rotation, mirroring, metadata, and every presented sequence frame.
- Complete the public AVIF format/API review so registered capabilities match implemented behavior.
- Remove or reject every valid in-scope AV1 syntax branch that remains silently ignored or unsupported.
Verified negative-path and frame-identifier gate evidence on 2026-08-31:
- A two-frame lossless frame-identifier sequence was generated and decoded with the clean official
libaom
maincheckout at observed revision441c439b9916474cac15d2822af47a9ad70674a8. Both decoded frames match the source Y, Cb, and Cr samples exactly. DecodeFrameIdentifiersMatchReferenceexecutes the production decoder through FeatureTestRunner, compares both native frames exactly, and proves the second frame is dependent with a changed current frame identifier. The current-frame, reference-delta, stale-slot, and refreshed-slot identifier logic was audited against the same currentmainsource.- The focused negative-path set passes 46 of 46 cases on net10.0 and 46 of 46 on net11.0, with zero failures or skips. It covers truncated palette entropy, malformed-following-OBU recovery, parser lifecycle failure, overflowing and invalid tile bounds, reference-slot ownership and transfer, constrained multi-group allocation, motion-field allocation failure unwinding, and frame identifiers.
- The established paused-stream cancellation suite now includes AVIF. It verifies cancellation at 0%, 30%, and 70% of both file and memory streams, plus pre-cancelled identification, on both targets.
- The completed checkpoint was committed as
7f0e08126b3354e8f1eb45886f0d572006ae27dewith author and committerJames Jackson-South <james_south@hotmail.com>.
Verified bounded-OBU checkpoint evidence on 2026-08-31:
- Audited
av1/decoder/obu.c,av1/decoder/decodeframe.c,av1/common/obu_util.c,av1/common/tile_common.c,aom/src/aom_integer.c, andaom_dsp/bitreader_buffer.cin the clean official libaommaincheckout. BothHEADandorigin/mainresolved to the observed revision441c439b9916474cac15d2822af47a9ad70674a8; this is verification evidence, not a pin. - The bounded container scanner and production OBU reader now agree with current libaom on ignored reserved header fields and the shared unsigned 32-bit LEB128 limit.
- Sequence-header validation now rejects undefined level indices, initial display delays above ten, frame identifiers above sixteen bits, zero timing units, the UVLC overflow sentinel, and invalid identity-matrix profile or subsampling combinations at the owning syntax boundary.
- Frame and tile parsing now rejects
show_existing_framein a combinedOBU_FRAME, the all-slots intra-only refresh mask, inner tile columns below current libaom's super-resolution-aware minimum, overflowing or out-of-bounds tile sizes, and empty final tile payloads. - The still-image writer now emits the required zero tile-bound-presence bit for a multi-tile combined
OBU_FRAME, matching current libaom's single-tile-group encoder path. ObuFrameHeaderTestsandObuFrameLifecycleTestscover the corrected syntax through the real bounded parser. The focused parser set passes 50 of 50 cases on net10.0.- The final focused production set passes 55 of 55 cases on net10.0 and 55 of 55 on net11.0, with zero failures or skips. It includes exact final-layer and selected-layer native planes, exact established reference-image presentation, constrained allocator ownership, malformed-following-OBU recovery, and FeatureTestRunner normal, AVX-512-disabled, AVX-disabled, and scalar execution.
- A fresh direct foreground current-main
aomdecrun decoded both progressive layers with one thread and row multithreading disabled. All 2,178 Y, U, and V samples match the retained YUV444-alpha reference; the alpha plane is excluded from the AV1 native-plane comparison. - The current-libaom production reference test and its established PNG were renamed together. The PNG
bytes remain unchanged at SHA-256
0758C17DC36E38AEE9F4389A335C2BF332AB91E4C79D7B0B22994FDDD0FD1605, both paths resolve todiff=lfs, and.gitattributeswas not edited. - Release source builds pass for net10.0 and net11.0 with zero warnings and zero errors. Roslynk reports zero compiler errors, and scoped production and test analyzer verification reports no changes.
- The completed checkpoint was committed as
243524c2c0b52a49d8d161fab806ab092cabe47cwith author and committerJames Jackson-South <james_south@hotmail.com>.
Verified partition, mode, segmentation, quantization, and transform checkpoint evidence on 2026-08-31:
- Audited partition traversal and chroma representability against
read_partitionand the subsampled plane-size rejection in current libaomav1/decoder/decodeframe.c; spatial segment-ID decoding and corruption handling againstread_segment_idinav1/decoder/decodemv.c; delta-Q syntax, resolution, arithmetic, and clamping againstread_delta_qindexandread_delta_q_paramsin the same file. - Audited selected and variable transform-size traversal against
read_tx_size,read_tx_size_vartx, and transform-block traversal inav1/decoder/decodeframe.c; coefficient syntax and arithmetic againstav1_read_coeffs_txbinav1/decoder/decodetxb.c; inverse quantization and transform application against currentav1/decoder/decodeframe.c,av1/common/idct.c, and the current libaom transform test oracle. The observed cleanHEADandorigin/mainrevision was441c439b9916474cac15d2822af47a9ad70674a8; this is verification evidence, not a pin. - Partition decoding now rejects an invalid partition subsize and a block size that cannot represent the current subsampled chroma plane. Spatial segmentation rejects decoded IDs above the active segment range. Focused tests exercise both current-libaom corruption boundaries through the production tile reader.
- Coefficient entropy decoding uses one allocator-owned maximum-size
Av1LevelBufferper tile reader. Each transform resets and clears only its active padded geometry, so no transform creates an allocation. Allocation tracking over all eight minimum- and maximum-quantizer frames proves exactly one coefficient scratch allocation per frame and exactly-once return after decoder disposal. - Palette index maps use one allocator-backed 32 KiB decoder-session owner with non-owning 128x128 luma
and chroma views. Each parsed superblock is reconstructed before either view is reused, and each block
clears only its transient
Buffer2DRegionafter prediction. The fixed session cost replaces the former full-frame maps without copies, fragmented memory groups, constructor rollback, or per-block allocations. The one-rent ownership regression and native palette reconstruction pass on net11.0, the complete HEIF/AV1 namespace passes 8,808 of 8,808 direct VSTest cases in Release, and the four-case palette set passes with exact native and presentation output, truncated-entropy rejection, and balanced exactly-once disposal. Av1BlockModeInfois value storage, removing the managed object allocation formerly created for every decoded coding block. ExplicitModeInfoIndexvalues preserve libaom's mode-info identity semantics at prediction-unit loop-filter edges, and the frame map now uses integer offsets so more than 65,535 decoded blocks cannot wrap its lookup identity.- Current official libaom reproduced the 39-frame all-intra reference and all four 8/10-bit minimum- and
maximum-quantizer references byte for byte. The production tests compare every native sample exactly,
cover every intra mode and seven selected transform types, execute SIMD and scalar paths through
FeatureTestRunner, and exercise the quantizer sequences under constrained tracked allocation. - Current official libaom decoded the 42-byte palette payload into the retained 1,089-byte YUV444
reference at SHA-256
E05F7C0DF06ECCF0E43869D1D7B03DAA1D635ACD26A766F8940899BE18D53251. The exact native test requires luma and chroma palette syntax. The established reference-output test uses the unchanged presentation PNG at SHA-2561148EBF6AA4B0F2D069D5E9B9605F6FB2A315E525F18016CDCAE23EFDD81DA84, whose renamed path still resolves todiff=lfs;.gitattributeswas not edited. - The exact final AV1 namespace passes 8,732 of 8,732 cases on net10.0 and 8,732 of 8,732 cases on net11.0, with zero failures or skips. Release source builds pass for net10.0 and net11.0 with zero warnings and zero errors. Roslynk reports zero compiler errors, and scoped analyzer verification reports no changes.
- The completed checkpoint was committed as
57a3f6668e39d0934e7b6b8d37a3dc2a5adc88f0with author and committerJames Jackson-South <james_south@hotmail.com>.
Verified retained-frame lifecycle checkpoint evidence on 2026-08-31:
- Audited primary-reference entropy selection, independent per-tile CDF starts, context-update-tile
publication, segmentation-map inheritance, reference-map refresh, and show-existing key-frame reset
against current libaom
av1/decoder/decodeframe.c,av1/decoder/decodemv.c,av1/decoder/decoder.c, andav1/common/entropymode.c. - Audited retained motion-vector cells, reference-side classification, projection source ordering,
projection limits, and reference-frame publication against
av1_copy_frame_mvs,av1_calculate_ref_frame_side,motion_field_projection, andav1_setup_motion_fieldin current libaom. Same-role primary-reference global-motion inheritance remains covered by the exact current-main global-warp fixture. The observed cleanHEADandorigin/mainrevision was441c439b9916474cac15d2822af47a9ad70674a8; this is verification evidence, not a pin. - Current official libaom decoded the retained
cdfupdate,mfmv,svc-L2T1,svc-L1T2, andsvc-L2T2streams with one thread, row threading disabled, and eight-bit output depth. Their generated Y4M files match the retained references byte for byte at SHA-2564FBFF73FF0DE2D9084DAE557D1D4BD677B0486516525BF4D327D2D795D5A7779,F7DB607694818C19E62FD9A27F53E1A3E2D00B72C39C0430C1B26399CC76777D,7A427631ECBF144F435AA4612F1201415FB1A9BCF9A67BA010AEF830B0C3AB81,4012DE2D4AFD095E7BB68EAE18B50B0674781BB4971CECABC0E5471E63373ED3, and1ABB981CFF76BA9557DA437B258D8A95FCA755DED8E3949D857E8388AB1D6AE3. - Existing allocation-tracking tests exercise initialization, retained-slot aliases, allocation-failure unwinding, presentation ownership, decoder-result ownership, repeated disposal, and final exactly-once return of reference frames, frame-owned motion fields, entropy snapshots, and segmentation maps.
- The focused Release checkpoint set passes 54 of 54 cases on net10.0 and 54 of 54 cases on net11.0, with zero failures or skips. It includes exact native CDF-update, motion-field, spatial-layer, temporal-layer, spatial-temporal-layer, progressive dependent-frame, and global-warp production paths, plus constrained allocator coverage.
- Release source builds pass for net10.0 and net11.0 with zero warnings and zero errors. Roslynk
reports zero compiler errors, scoped analyzer verification reports no changes,
git diff --checkpasses, and.gitattributesis unchanged.
Final decoder allocation, lifetime, precision, architecture, and test-validity audit evidence on 2026-09-01:
- Refreshed the official libaom remote and audited against observed
origin/main976867526367f571a1c09b994066af8364aed781. The intervening external-rate-controller commit does not changeav1/decoder,av1/common,aom_dsp, or the AV1 decoder build definition. - CDEF now uses one bounded 64x64-unit bordered source workspace, two preserved top-row slots per plane, preserved left columns, and unit-local direction and variance storage. This replaces the frame-wide source copy and frame-wide direction maps while retaining libaom's unit traversal and cross-plane luma-direction lifetime.
- Loop restoration now retains the required immutable source and separate destination, but stores the
full destination in native sample width. Eight-bit filtering narrows only bounded unit output after
clipping, while high-bit-depth filtering writes directly to the native
ushortdestination. - Reference-to-presentation copying now copies visible native rows only. Padding remains destination owned, and the ownership tests mutate a copied visible sample rather than unrelated padding.
- The remaining decoder allocations and copies are either bounded scratch or required ownership boundaries. Frame planes enforce their contiguous single-span invariant before allocation; palette, transform, film-grain, super-resolution, color-conversion, and alpha workspaces remain bounded and allocator owned. No per-block managed allocation remains in reconstruction.
- Block reconstruction now uses one exact-size signed-short owner for inverse quantization, inverse transform, compound prediction, convolution, and chroma-from-luma scratch. Even-length slices provide the integer workspaces without another rent. Monochrome reserves no chroma coefficients, and 4:2:0, 4:2:2, and 4:4:4 reserve two symmetric chroma planes at their coded subsampling. This replaces three constructor rents and their catch-all rollback path; exact allocation length, coefficient span length, and exactly-once return pass for all four layouts, with 549 adjacent reconstruction tests passing direct net11 VSTest in Release.
- Valid unsupported tile-list syntax is rejected explicitly. Reserved and metadata OBUs are consumed only after bounded framing and trailing-bit validation. Eight-, ten-, and twelve-bit reconstruction, presentation, alpha, restoration, and film-grain paths retain native precision.
- Predictor traversal remains split into semantic readonly operator families. The planar sample
adapter and transform-block context are value types, and Release construction sites use
defaultwithout null-forgiving suppression. - The net11.0 Release test project builds with zero errors. Roslynk reports zero compiler errors,
git diff --checkpasses, and.gitattributesis unchanged. - Visual Studio 18.9 VSTest ran the complete
Formats.Heif.Av1namespace with collection parallelism disabled and stop-on-failure enabled: 8,746 of 8,746 cases passed. The touchedHeifDecoderTestsandHeifSequenceParserTestsadd 104 of 104 passing integration cases. Focused CDEF, restoration, film-grain, copy-ownership, and reference-isolation runs also pass 15 of 15 cases. No test-host crash or Windows application-error dialog occurred.
Final decoder stream, presentation, and public-registration evidence on 2026-09-01:
- Real AV1 still-item and timed-sequence files decode identically from a file stream, memory stream, non-seekable stream, and a seekable stream limited to three bytes per read. All eight stream rows pass through public format detection and production decoding, comparing every presented frame exactly.
- A two-frame production sequence applies a centered clean-aperture crop, counter-clockwise rotation, mirroring, pixel-aspect-ratio metadata, and CICP metadata to every frame. The complete five-frame real auxiliary-alpha sequence composes non-opaque alpha and retains timing, Exif, and XMP for every frame.
- The fixed-header detector accepts both compact and extended-size leading file-type boxes. Default configuration registers the implemented HEIF decoder and detector but no longer advertises the incomplete HEIF encoder.
- Visual Studio 18.9 VSTest, serialized with stop-on-failure enabled, passes the 12 of 12 new
stream/presentation/registration cases and the complete current
HeifDecoderTestsplusHeifSequenceParserTestsset with the registration contract: 115 of 115. The final explicit no-encoder registration assertion passes 1 of 1 after its final edit. - The net11.0 Release test project builds with zero errors, Roslynk reports zero compiler errors,
git diff --checkpasses, and.gitattributesis unchanged. Every VSTest invocation returned normally with no surviving test host and no Windows application-error dialog.
SIMD traversal consistency evidence on 2026-09-02:
- The shared
Numericsvector-count helpers now cover same-lane spans and all fixed hardware widths. AV1 decoder and current encoder hot paths use those helpers for complete-vector traversal instead of repeating local modulo or last-vector calculations. Reverse-source indexing and algorithm-specific partial-output groups remain explicit because they are not vector-count calculations. - Forward quantization, palette prediction, and scaled inter prediction construct width-specific SIMD constants only when at least one vector batch will execute. Narrower dispatch tiers consume only the remainder left by wider tiers before the scalar tail.
- The net11.0 Release production assembly builds with zero warnings and zero errors. The test project
builds with zero errors while retaining the existing repository warning set. Roslynk reports zero
compiler errors,
git diff --checkpasses, and.gitattributesis unchanged. Foreground VSTest passes 63 of 63 focused quantizer, forward-transform, CDEF, restoration, palette, intra, inter, film-grain, and super-resolution cases.
Decoder exit gate:
- Every supported native format and AV1 tool has exact current-main libaom production-path evidence.
- Every supported presentation behavior has established reference-image evidence at the correct output precision.
- No decoder path relies on a native codec, copied plane, per-block allocation, or contiguous memory-group accident.
- All allocator ownership is deterministic and exactly once.
- Full focused Release verification is recorded with no false coverage claims.
AV1 encoder implementation
Writer primitives are not an encoder. The public encoder remains incomplete until it produces independently decodable AV1 payloads and AVIF containers for every exposed option.
5. Define and enforce the encoder contract
- Use official libaom
mainata40ed1ea9e4ecc3df58a5bccb76623f2c94ae727as the encoder syntax, probability-model, transform, quantization, filtering, and bitstream reference. - Use the existing PNG, TIFF, and JPEG encoders as the ImageSharp architecture reference: generic
Image<TPixel>input, encoder options taking precedence over converted format metadata and codec defaults, allocator-owned temporary storage, and deterministic disposal. - Treat source pixel type, source alpha representation, and decoded source bit depth as conversion inputs, never as output-eligibility checks. Do not pre-scan pixels before encoding.
- Resolve output configuration once from explicit encoder options, converted
HeifMetadata, and AV1 defaults in that order. Sanitize only combinations that cannot describe a legal requested output, and never write resolved values back to source metadata. - Finalize observable options for quality, effort, lossless mode, bit depth, chroma subsampling, alpha quality, metadata, and bounded sequences.
- Preserve high-bit-depth source precision through 16-bit RGB and native 10/12-bit component planes.
- Reject only genuinely unsupported output combinations at the public boundary before writing output.
- Register only capabilities that the completed encoder proves.
Encoder data-flow contract:
- Resolve immutable frame and sequence output settings before allocating codec state.
- Convert each generic
ImageFrame<TPixel>once throughPixelOperations<TPixel>and the SIMD-first HEIF planar converter into native 8, 10, or 12-bit planes. Alpha is encoded as an auxiliary image when requested by the resolved output contract; it is not discarded through a source scan. - Reuse allocator-owned plane, row, block, transform, quantization, entropy, and reconstruction workspaces for the complete frame. No active path may allocate per row, block, transform, scanline, or SIMD tail.
- Analyze and encode tiles directly from those planes, retaining reconstructed reference frames only for the bounded sequence lifetime.
- Build OBU headers in bounded allocator-backed scratch and stream entropy-coded tile owners and container extents directly. Every ownership transfer is explicit, every owner is disposed exactly once, and no
ToArrayor file-sized copy crosses a layer boundary. - Iterate image frames using ImageSharp frame metadata and format-connecting metadata. Root-frame-only behavior is permitted only for an explicitly static output contract.
Encoder verification contract:
- Exercise source pixel formats independently from requested AV1 bit depth, chroma subsampling, alpha, and lossless/lossy mode.
- Run every SIMD operator through FeatureTestRunner at Vector512, Vector256, Vector128, and scalar tiers against an independent scalar oracle shaped from the same libaom revision.
- Cover discontiguous allocator buffers, constrained memory groups, cancellation, non-seekable output, multiple extents, auxiliary alpha, and bounded sequences.
- Validate produced AV1 payloads with current-main libaom and compare native planes before using ImageSharp self-decode as supplemental container coverage.
6. Build the complete AV1 frame encoder
-
[~] SIMD-first RGB-to-native-plane conversion now feeds eight-bit and high-bit-depth bordered AV1 source frames directly, preserving ImageSharp's arbitrary packed-pixel input contract without an intermediate full-frame native-plane copy.
-
[~] Auxiliary-alpha encoding now follows the same packed-pixel conversion boundary without scanning pixel contents or cloning the image. Source alpha is converted through ImageSharp's 16-bit pixel contract, deinterleaved with descending Vector512, Vector256, Vector128, and scalar traversal through the shared vector-count helpers, then scaled and rounded once by the existing native-sample writer directly into the final bordered monochrome source frame. One operation-wide allocator owner provides the packed and planar row views; there is no frame-sized alpha staging allocation or second owner. Exact 12-bit precision, physical border extension, the single 12-bytes-per-pixel row rent, and balanced return pass through the production converter. The complete 47-case frame-encoder set passes direct foreground net11 Release VSTest, and current-main
aomdecata40ed1ea9e4ecc3df58a5bccb76623f2c94ae727accepts the generated 8-, 10-, and 12-bit monochrome payloads. AVIF auxiliary item properties, references, and public activation remain open. -
[~] Forward transform families, transform workspace, and an allocation-free DC intra block boundary exist locally. For eight-bit and high-bit-depth samples, the composed boundary now follows current libaom's encoder order: predict into the reconstruction plane, subtract prediction from source, transform, quantize into separate qcoeff and dqcoeff storage, retain EOB and transform type, and inverse-transform only when EOB is nonzero so later blocks consume decoder-identical references. Prediction and subtraction retain their SIMD-first operators, independent source and reconstruction strides are preserved, and no frame-sized or per-block buffer is introduced. The block boundary consumes the real bordered encoder-plane regions and indexes their one-segment owner directly; this preserves physical row strides without a row copy and avoids the per-call enumerator allocation exposed by the initial array-only test. One reusable 61 KiB allocator owner supplies tightly packed residual, aligned transform-coefficient, dequantized-coefficient, and transform scratch spans across transform blocks; quantized coefficients write directly to the retained frame coefficient owner instead of being duplicated. A fixed 8x8 DC-intra superblock baseline now traverses the same recursive preorder and frame-edge pruning as the tile writer, gathers left references into that reusable block workspace, writes luma and chroma coefficient-owner slices in the writer's exact consumption order, and updates the caller-owned reconstruction planes for subsequent predictions. Stage-by-stage scalar-oracle, physical-border, retained-syntax, superblock-to-writer synchronization, high-bit-depth precision, and steady-state zero-allocation coverage passes 8 of 8 through direct net11 VSTest in Release. This is a legal fixed baseline, not complete partition or mode analysis.
-
[~] A production single-tile all-intra writer now walks raster superblocks, analyzes each immediately before entropy coding, reuses one decision workspace and one block workspace, and retains decoder-identical reconstructed references across the tile. Its closed byte and high-bit-depth operators feed the existing superblock boundary without runtime sample-type checks. A byte-exact test compares this composed path with an explicit superblock-then-tile-writer oracle, so producer and writer traversal or coefficient-area drift cannot pass unnoticed. A separate clipped 2x2-superblock regression proves global raster indexing by requiring all four coefficient segments and the bottom-right reconstruction to be populated. Multi-tile ownership and the complete frame/OBU operation remain.
-
[~] A non-owning encoder-frame view now separates visible conversion regions from coded regions and performs complete left, top, right, bottom, and corner extension across each bordered plane. Current libaom uses 8-sample-aligned coded dimensions, a 32-sample-aligned luma stride with chroma stride derived from it, and a 64-pixel luma border for non-resized all-intra encoding. One operation-ready frame owner now rents the aligned Y, U, and V storage contiguously, exposes non-owning
Buffer2Dplane views, and returns the rent exactly once. A 4K 4:2:0 frame occupies about 13.0 MiB at 8-bit or 26.0 MiB at 10/12-bit; source and reconstruction therefore remain distinct frame owners rather than adding a full-frame copy. The corrected tests use this real ownership path and verify the exact 54 KiB 64x64 4:2:0 rent. The frame-encoder operation now instantiates matching source and reconstruction owners with ordinaryusinglifetimes and converts packed pixels directly into the source owner before extension. -
[~] Temporal delimiter, sequence header, frame header, combined-frame tile-group writing, and an internal reduced-still-picture frame operation now exist locally. The remaining required metadata, padding, multi-tile, option, and public encoder paths are not complete.
-
[~] Implement superblock and partition analysis for every permitted block size and partition. The current baseline deliberately splits every in-frame node to 8x8 blocks and records decisions in current-libaom writer preorder; block-size selection and non-split partition analysis remain.
-
[~] Implement intra mode search, palette, filter intra, chroma-from-luma, and intra-block copy decisions. Live luma search now covers all 13 zero-angle base modes and all six nonzero adjustments for each of the eight directional modes. Joint spatial chroma search covers the same 61 candidates, combines both chroma planes in one rate-distortion decision, and preserves the winning shared angle adjustment. Chroma-from-luma now searches the complete signed alpha alphabet from reconstructed luma and retains its joint U/V syntax. Filter-intra now searches all five predictors after ordinary luma modes. Palette entropy, retained state, production syntax, exhaustive luma and paired chroma palette selection, adaptive screen-content activation, and joint intra-block-copy mode selection are complete.
-
Implement inter mode search for bounded sequences, including reference selection and the decoder-supported inter tools.
-
[~] Current-libaom
av1_quantize_fp_no_qmatrixarithmetic is implemented as a closed generic forward-quantizer family with Vector512, Vector256, Vector128, and scalar paths, raster-order output, coded 64-point coefficient limits, and scan-order EOB selection. Transform search, coefficient optimization, and lossless behavior remain. -
[~] Implement real rate-distortion selection and make quality and effort change work, size, and output quality. The complete luma and joint chroma candidate sets, including chroma-from-luma, filter-intra, palette, and intra-block copy, now perform live rate-distortion selection; quality mapping, effort tiers above the current uniform-transform search ceiling, and the remaining searches are not implemented.
-
[~] Frame effort now progressively expands the available current search: zero is DC-only, one adds every zero-angle spatial mode, two adds every legal directional adjustment, three adds transform refinement, four adds filter-intra and chroma-from-luma, and five adds adaptive palette and intra-block-copy analysis. Lower tiers do not signal unavailable sequence or frame tools, and tiers below five skip the whole-frame screen-content scan. Effort six enables
TX_MODE_SELECTand compares the winning ordinary spatial luma mode as one 8x8 transform against four raster-ordered 4x4 transforms. Every 4x4 transform searches all legal transform types with live coefficient contexts and reconstructed intra references. The search reuses the aligned block workspace, preserves only improving candidates, and performs no per-block or per-transform rent. Non-skipped intra-block copy writes and costs the current-libaom unsplit variable-transform root; skipped intra-block copy emits no transform-partition symbol. Values seven through ten currently share the effort-six ceiling. Transform-size integration for filter-intra and palette, broader joint mode/transform refinement, partition search, and pruning remain. Ten decoder-visible production cases verify emitted flags, mode restrictions, real 4x4 selection, intra-block-copy syntax, and successful decode. The exact net11 Release test-project build remains at 1,005 baseline warnings and zero errors; the clean complete non-HEVC HEIF/AV1 namespace passes 9,294 of 9,294 through direct foreground VSTest. Current-mainaomdecata40ed1ea9e4ecc3df58a5bccb76623f2c94ae727accepts both new effort-six payloads in addition to the previously verified set; their decoded-frame MD5 values are2dd1cbe449fe2d0471dc2c15c50acb69for spatial 4x4 transform selection and677435e5af39c930af1178f91c34af6afor intra-block copy. Roslynk reports zero compiler errors and no analyzer diagnostics in the touched files. -
[~] Encoder rate accounting converts the entropy writer's live inverse cumulative distributions into current-libaom fixed-point symbol costs without allocating or duplicating probability state. Read-only luma-mode, directional-delta, filter-intra, chroma-mode, block-skip, transform-size, transform-block-skip, and complete transform-coefficient queries share the exact distributions mutated by the subsequent entropy write. Complete coefficient costing follows current libaom's optimized shape: it returns immediately for an empty transform, uses the EOB-specific base-range context, fuses magnitude, sign, base-range, and Golomb accounting into one reverse traversal, and combines repeated full base-range chunks instead of replaying each emitted symbol. Tile-lifetime level and context scratch is reused, the one-coefficient path neither clears nor initializes the forward-neighbor level map, and steady-state queries allocate nothing. Transform-size writing and costing share one subdivision-depth calculation, while shared closed symbol operations keep the writer and cost mappings for transform skip, transform type, and EOB syntax identical without forcing the estimator through the writer's slower two-pass coefficient traversal. The current-libaom fixed-point RD combiner preserves 64-bit distortion and rounds the weighted 1/512-bit rate at the required boundary. Its key-frame multiplier follows libaom's squared DC-quantizer formula and exact 10/12-bit normalization. Live final-block selection evaluates all 61 legal 8x8 luma candidates: the 13 zero-angle base modes in current-libaom order, followed by six nonzero adjustments for each directional mode. Joint chroma selection evaluates the equivalent 61 spatial candidates, combines U and V distortion plus coefficient rate, and charges one live chroma-mode and shared-angle symbol over the actual subsampled 4x4, 4x8, or 8x8 geometry. Chroma-from-luma subsamples the reconstructed luma block once into fixed-stride Q3 stack scratch, subtracts the rounded mean, evaluates all 33 signed alpha values independently for each plane with complete transform RD, and combines the cached plane results across all 1,088 valid joint pairs with one live sign cost and the conditional U/V magnitude costs. This is the allocation-free equivalent of current libaom's exhaustive 33-value path: it requires 66 evaluation transforms rather than transforming every joint pair, preserves DC-before-CfL-before-spatial tie order, and fixes the implicit chroma transform to DCT-DCT. Filter-intra follows ordinary luma candidates, searches all five predictors in syntax order, and evaluates every legal transform while reusing one prepared prediction and source residual per filter mode. Every candidate includes its live mode, angle, filter mode, alpha, and coefficient rate plus normalized pixel-domain distortion. Each prepared reference edge retains the common-corner prefix and twice the transform dimension required by directional prediction. A shared encoder/decoder availability calculation selects reconstructed top-right and bottom-left extensions according to tile, frame, superblock, and block reconstruction order; unavailable extensions repeat the nearest coded endpoint. Missing top or left edges retain current libaom's perpendicular-sample and bit-depth-midpoint rules. Directional prediction applies the AV1 three-degree adjustment step and reuses transform workspace for zone-three transposition before the transform overwrites it, keeping candidate evaluation allocation-free. The winning luma and chroma signed adjustments are retained in the packed final-block state consumed by the tile writer. The tile writer invokes these reusable workspace-backed selectors after mapping current neighbors and immediately before writing each block, so later decisions see reconstructed samples, coefficient contexts, and CDF updates from every preceding block. Block skip is read only after the callback has combined every coded plane. Luma and chroma candidate scratch is partitioned from the encoder's single aligned reusable block workspace; transform-size search uses that owner for four retained 4x4 transform states, local coefficient contexts, and the compact trial reconstruction needed to preserve the best result. No candidate path rents a buffer per block or per transform. Only a newly winning candidate is copied into retained frame storage. Production fixtures force every luma base predictor, both extreme adjustments in all three directional zones, available top-right and bottom-left extensions, high-bit-depth adjustment propagation, exact signed luma and chroma angle-rate terms, joint U/V decisions, packed chroma state, and 4:2:0, 4:2:2, and 4:4:4 transform geometry. The CfL fixtures derive target chroma from a pilot production encode's actual reconstructed luma through an independent scalar Q3 oracle and prove exact positive/negative alpha syntax plus zero-residual DCT-DCT reconstruction for all three subsampling geometries at 8, 10, and 12 bits. The stable fixed-DC traversal comparison uses neutral samples for which both the baseline and live search are contractually DC and skipped, instead of relying on textured content to happen to select the baseline mode. Luma palette selection now evaluates dominant-color and one-dimensional K-means candidates for every legal size, snaps near-cache colors with the reference threshold and tie order, removes duplicate snapped colors, extends boundary maps from active samples, and performs complete transform rate-distortion search. Ordinary DC and filter-intra candidates pay the palette-disabled symbol whenever screen-content syntax is enabled. The exact net11 Release rebuild reports 1,992 test-project warnings and zero errors, all 58 intra-superblock cases pass, all 8,935 AVIF cases pass, and all 230 HEIF cases pass. Remaining mode decision work includes transform-size coverage for filter-intra and palette, broader joint mode/transform refinement, partition search, and effort-dependent pruning. Non-empty intra blocks deliberately remain non-skipped, matching current libaom; later inter mode selection owns its distinct skip-transform RD decision.
-
[~] The tile writer now publishes one packed coefficient context per covered 4x4 edge unit and derives luma/chroma skip plus DC-sign contexts from the complete transform edges using current-libaom units. Partition, transform, and coefficient neighbor state retains only the above and left context regions used by current libaom; the unused third top-left region, its granularity state, and its unused sentinel are removed. One picture owner now packs segmentation plus every tile's partition, luma, chroma, and transform edges into one clean byte allocation with typed non-owning views; together with the separately typed packed mode-information owner, the complete picture state uses two allocator rents rather than seven. Exact aligned lengths, clean initialization, and balanced exactly-once returns are covered in Release. Multi-tile payload ownership and verified CDF update behavior remain.
-
[~] Encoder mode information now uses a frame-owned integer alias grid over a packed 8-byte value allocation, matching current libaom's
mi_grid_baseandmi_allocrelationship without a managed object or reference per 4x4 entry. The visible dimensions are aligned to eight luma samples, the grid stride and allocated row count are aligned to 32 mode-information units, and optional 8x8 allocation granularity reduces the value store in both dimensions exactly as current libaom does. One clean ImageSharp byte owner contains both independently typed regions, reducing libaom's two allocation lifetimes to one without a copy. At 4K, the 4x4 layout occupies about 6.0 MiB in total; the 8x8 layout occupies about 3.0 MiB. Exact geometry, clean allocation, typed lengths, aligned mapping, untouched row padding, and exactly-once return pass 4 of 4 direct net11 VSTest cases in Release. Every coded 4x4 cell covered by square, rectangular, or clipped edge blocks maps to its owning allocation entry before context-dependent symbols are written. Packed syntax, relative neighbor lookup, full block mapping, writer traversal, entropy, and OBU coverage pass 1,947 of 1,947 direct net11 VSTest cases in Release; complete mode decision still remains. -
[~] The final-block decision workspace uses one reusable 8.3 KiB ImageSharp allocator owner. It contains 1,024 explicitly packed 8-byte final-block entries and the 341 preorder partition bytes required by a complete 128x128-through-8x8 quadtree, replacing separate managed arrays. Palette colors now have their own current-block value and are copied only to the picture edges that later blocks can reference, so enabling palette mode does not add 50 bytes to every final-block entry. Construction and the explicit per-superblock reset initialize every syntax field, including the nonzero sentinel that disables filter-intra prediction; pooled quantizer, prediction, partition, and current-palette bytes cannot leak into the next decision pass. Complete mode decision still remains.
-
[~] Finalized transform coefficients and packed EOB/type state now use raster-ordered, per-superblock plane segments matching current libaom's coefficient-pool geometry. One ImageSharp allocator owner replaces libaom's separate coefficient, EOB, and entropy-context allocations while preserving the full 1024 luma and 256-per-chroma 4x4 state capacity of a 128x128 4:2:0 superblock. The fixed 8x8 DC-intra traversal populates the owner's quantized coefficient and state slices while updating the caller-owned reconstruction plane directly, and a real tile-writer integration check proves that both sides consume identical luma and chroma areas. Complete mode decision still remains.
-
[~] Tile partition writing now follows current libaom's recursive
write_modes_sbpreorder traversal andupdate_ext_partition_contextedge updates directly. Bottom-edge blocks use the horizontal-alike partition CDF and right-edge blocks use the vertical-alike CDF; byte-exact regressions cover both paths after the previous calls were found reversed. Lossless chroma-from-luma availability now uses the subsampled plane block size shared with the decoder instead of the lossy 32x32 limit, preserving the correct UV-mode alphabet for each segment. The obsolete SVT-derived global geometry catalog and its unimplemented lookup are removed; transform geometry is derived in libaom's bounded 64x64 residual order, fixed intra transform-size symbols use the reference depth and neighbor contexts, and each derived transform size is persisted to the frame-owned mode information before the entropy snapshot and coefficient traversal consume it. Frame-edge and segmentation syntax use mode-information units, and 128x128 CDEF units use libaom's 0-to-3 indexing and first-block strength ownership. The focused transform-state regression passes 3 of 3 direct net11 VSTest cases in Release. Writer, entropy, and OBU coverage passes 1,957 of 1,957 direct net11 VSTest cases in Release, with 20 of 20 focused encoder and decoder chroma-from-luma cases. Partition and mode analysis still need to populate these retained decisions; variable inter-transform syntax remains part of later inter-frame support. -
Implement legal deblocking, CDEF, restoration, super-resolution, and film-grain signaling decisions.
-
[~] The coefficient symbol encoder now reuses tile-lifetime level and context workspaces instead of allocating per transform, defers both coefficient rents until the first nonzero transform block, and disposes all tile scratch independently from the detached encoded bytes. Its range coder matches current libaom's 64-bit coding window, bulk big-endian byte flush, and backward carry propagation while using one byte of allocator scratch per estimated output byte instead of the former 16-bit pre-carry storage. The single-tile production path finalizes in that existing allocation and transfers its owner plus the used byte length, removing the former second rent and full-tile copy; exact-length test callers retain the original overload. The ownership regression proves that writer disposal cannot return transferred storage and that the caller returns the original allocation exactly once. The internal frame operation now passes that payload directly to the OBU writer; AVIF container integration and public activation remain.
-
[~] The planar conversion, DC intra prediction, residual construction, forward transform, and forward quantizer use descending SIMD dispatch: Vector512, Vector256, Vector128, then scalar. Residual construction matches current libaom's exact source-minus-prediction arithmetic for 8-bit and high-bit-depth planes, preserves independent row strides and unaligned starts, and writes directly into caller-owned signed-short storage without allocation. Candidate distortion reuses that residual workspace and widens signed 12-bit lanes before vector squaring, accumulating exact full-block SSE in 64-bit scalar storage. The composed block path delegates arithmetic to those closed operators and adds no allocation. Apply the same rule to every later hot-path family.
-
[~] Residual tests verify misaligned planes, independent source, prediction, and destination strides, SIMD remainders, untouched padding, 8-bit, 10-bit, and 12-bit precision, every operator width independently of host acceleration, the scalar fallback, and zero per-transform allocations.
-
[~] The unused coefficient-shape transform facade and its unimplemented N2, N4, and DC-only branches are removed. Finalized block encoding now follows the complete-transform path that current libaom uses before fast quantization; later rate-distortion search may add proven coefficient optimization without exposing inactive runtime throws.
-
[~] Forward-quantizer FeatureTestRunner and zero-allocation tests compare every hardware tier with an independent scan-order scalar oracle shaped from current-main libaom. Both passed direct net11 VSTest in Release.
-
[~] The combined-frame writer now completes the byte-counted uncompressed frame header before starting the optional multi-tile tile-group flag, matching current libaom's separate frame-header and tile-group writers. A non-uniform two-tile round trip verifies the explicit boundaries, both tile payloads, and complete stream consumption through direct net11 VSTest in Release.
-
[~] The first internal frame-to-OBU operation encodes 8-, 10-, and 12-bit monochrome reduced still pictures through the production tile writer and production decoder. Coefficient context initialization now stores
min(abs(level), 127), matching current libaom; the previous signed clamp converted every negative transform coefficient to zero and selected invalid nonzero-map distributions. Signed dense and sparse entropy round trips, direct level-buffer saturation coverage, and eight constant/gradient frame cases pass 52 of 52 direct net11 VSTest cases in Release. Current-mainaomdecaccepts all eight emitted payloads. After winner-mode transform refinement, their decoded-frame MD5 values ared09ea148582b9c93fa78e59426193bbc(16x16 8-bit constant),b83eedd5a84428f0120130253b30bdaa(16x16 8-bit gradient),f949f7422913e83dff07ee5e0a5087d3(8x8 8-bit constant),ae7233a94558978934469dcc4da764dd(8x8 8-bit gradient),09223b227f3abc3134d0a3ea15f70c0a(8x8 10-bit constant),539aab0e6e14bcaec271febfa8e25444(8x8 10-bit gradient),73117a8fc102e5d028f82444fc4d15ab(8x8 12-bit constant), and6936a2b62d7220dfb12f3763bb49965d(8x8 12-bit gradient). This is an independently decodable baseline, not completion evidence for chroma, alpha, options, containers, or the public encoder. -
The exact net11 Release rebuild completed at the established 1,005-warning repository baseline with zero errors. The complete HEIF/AV1 namespace passes 8,838 of 8,838 direct VSTest cases with zero failures or skips. Roslynk reports zero compiler errors and no diagnostics in the five changed C# files;
git diff --checkpasses and.gitattributesis unchanged. -
[~] The same internal frame operation now produces 4:2:0, 4:2:2, and 4:4:4 payloads at 8, 10, and 12 bits. Twenty-one color cases cover constant and spatially varying input at aligned dimensions plus odd 13x11 visible dimensions for every chroma geometry. The production decoder consumes every payload, the decoded output retains non-neutral chroma, and current-main
aomdecaccepts all 29 monochrome and color outputs. After live spatial chroma mode selection, implicit chroma-transform correction, winner-mode luma-transform refinement, exhaustive chroma-from-luma alpha selection, and filter-intra search, the odd-dimension decoded-frame MD5 values are9985f05790d2c9f5f28723ef86d5b89b(4:2:0),2ba2f1d0fcfef60394a5175553c7cb8b(4:2:2), and6aa7a2ed0dbf76ad2ec0c222585272d0(4:4:4). This proves legal current-libaom payload syntax across native plane geometries; it does not yet prove target quality or native-plane equality with an independently encoded reference. -
Spatial chroma candidates now use the implicit transform derived from the selected UV mode and active transform set, matching current libaom's
intra_mode_to_tx_typeandav1_get_tx_typebehavior. The same shared derivation is consumed by the decoder, so coefficient scan order, entropy contexts, inverse reconstruction, and encoder rate estimates cannot drift between the two paths. The previous DCT-DCT candidate transform could produce syntactically accepted streams whose non-DC chroma coefficients were interpreted under a different implicit transform. Six production mode-decision cases retain nonzero U and V coefficients and assert the selected transform state across 4:2:0, 4:2:2, and 4:4:4; fifteen exact mapping cases cover every intra mode, reduced sets, and the 32x32 DCT-only fallback. The focused contract passes 21 of 21 direct net11 VSTest cases, the complete HEIF/AV1 namespace passes 8,947 of 8,947, the exact Release rebuild remains at 1,005 warnings and zero errors, and current-mainaomdecaccepts all 29 regenerated payloads. -
[~] Luma mode selection now evaluates each of its 61 mode-and-angle candidates with the mode-derived default transform used by current libaom's fast intra path. It then refines only the winning mode across all seven transform types permitted by the 8x8 intra set in transform-enum order. This removes the fixed DCT-DCT limitation while avoiding a 61-by-7 expansion; each trial includes live transform-type and coefficient rate, reconstructed pixel-domain distortion, and the existing aligned reusable block workspace. Eighteen exact-prediction production cases prove DCT-DCT wins equal-cost ties in reference order even when the first pass used a different default, while the 72x72 textured traversal proves a non-DCT transform with nonzero coefficients reaches retained syntax. Current-main
aomdecaccepts all 29 regenerated payloads. Special-mode transform-size coverage, full partition search, and broader effort-dependent joint mode/transform search remain. -
Chroma-from-luma mode decision now reuses the decoder's SIMD-first 4:2:0, 4:2:2, and 4:4:4 reconstructed-luma preparation and prediction kernels for both byte and high-bit-depth encoder operators. The constant DC predictor for each chroma plane is computed once and its sample refills every alpha candidate, matching libaom's per-plane DC cache instead of rebuilding the same edge average 33 times. Each block uses 512 bytes of fixed stack scratch for the maximum 8-row predictor surface plus 792 bytes for complete U/V rate and distortion tables; no allocator owner, managed object, frame copy, or persistent buffer was added. Live probability costs exactly mirror current libaom's joint-sign ownership and conditional magnitude symbols. Nine production cases independently derive exact CfL targets from decoder-visible reconstructed luma at 8, 10, and 12 bits, and three entropy cases cover two nonzero signs plus each single-zero-plane form. The exact net11 Release rebuild remains at 1,005 warnings and zero errors, all 8,959 HEIF/AV1 tests pass, and current-main
aomdecata40ed1ea9e4ecc3df58a5bccb76623f2c94ae727accepts all 29 regenerated payloads. -
Filter-intra mode decision now runs after ordinary luma modes in current-libaom order, evaluates all five recursive predictors, and refines each predictor across every legal 8x8 transform in transform-enum order. Strictly-better replacement preserves ordinary-mode and filter-mode tie order. Each filter prediction and its source residual are prepared once and reused across transform candidates, avoiding repeated recursive prediction while retaining SIMD-first predictor and subtraction operators. The stack cost is 192 bytes for eight-bit samples or 256 bytes for high-bit-depth samples; no allocator owner or managed buffer was added. Fifteen production cases force every filter mode at 8, 10, and 12 bits and prove retained filter syntax, zero-residual reconstruction, and the DCT-DCT equal-cost transform tie. The decoded-frame MD5 values selected by this checkpoint are
d7d68803763b95827483f14515281d3afor the 8x8 10-bit gradient,3f7e34d44c65d7797ad26b5cd4c35bf4for the 8x8 12-bit gradient, and9985f05790d2c9f5f28723ef86d5b89b,2ba2f1d0fcfef60394a5175553c7cb8b, and6aa7a2ed0dbf76ad2ec0c222585272d0for the odd 4:2:0, 4:2:2, and 4:4:4 gradients. The exact net11 Release rebuild remains at 1,005 warnings and zero errors, 18 focused filter-intra, predictor-reference, syntax-cost, and allocation cases pass, all 8,974 HEIF/AV1 tests pass, and current-mainaomdecata40ed1ea9e4ecc3df58a5bccb76623f2c94ae727accepts all 29 regenerated payloads. -
Empty-transform block skip now compares the complete live rate of the two decoder-identical syntax choices after luma and every coded chroma plane have been selected. Current libaom forces all-intra blocks to non-skip; this encoder retains that behavior for every non-empty block and for equal-cost empty blocks, but emits block skip when its adapted context cost is strictly lower than non-skip plus all empty-transform coefficient costs. Costing and writing share the same above-and-left skip-context calculation, and the coefficient estimator returns after the transform-block-skip symbol without reading coefficient storage. This adds no allocation, copy, or persistent state. A focused adapted-CDF regression proves both outcomes through the production decision helper, the two production all-zero fixtures still prove the default real block path, the exact net11 Release rebuild remains at 1,005 warnings and zero errors, all 8,975 HEIF/AV1 tests pass, and current-main
aomdecata40ed1ea9e4ecc3df58a5bccb76623f2c94ae727accepts all 29 regenerated payloads. -
[~] Palette entropy coding now mirrors current libaom's adaptive luma-mode, chroma-mode, palette-size, and spatial color-index distributions, together with its truncated-binary uniform code used by palette colors. The complete mutable palette probability graph is created once on first palette search or write, so the current palette-disabled frame path retains zero palette allocations. Three focused regressions cover every legal 2-through-8 color alphabet and every defined mode, size, and color-index context; all 1,928 entropy cases and all 8,978 HEIF/AV1 cases pass direct net11 Release VSTest. The exact Release rebuild remains at 1,005 warnings and zero errors. This checkpoint adds the exact entropy foundation only: palette candidate generation, retained color and index storage, mode decision, map tokenization, and production syntax remain incomplete, and no generated payload changed.
-
[~] Luma and chroma palette-color coding now matches current libaom's neighbor-cache flags, sorted delta representation, wrapped V-plane deltas, strict delta-versus-raw V selection, and fixed-point color-rate model at 8, 10, and 12 bits. Encoder costing and emission use only fixed stack spans, including explicitly initialized cache-membership state, and steady-state color costing allocates zero managed bytes. The decoder consumes the same bounded color-syntax primitive after the tile reader derives its neighbor cache, removing duplicated color parsing without changing retained palette ownership. Nine focused syntax, exact palette decode, constrained-allocation, truncation, presentation, and allocation cases pass; all 1,933 entropy cases and all 8,983 HEIF/AV1 cases pass direct net11 Release VSTest. The exact Release rebuild remains at 1,005 warnings and zero errors. Retained encoder palette colors, neighbor caches, color-index maps, candidate generation, and production palette selection remain incomplete, and the compact 8-byte frame mode entries were not enlarged.
-
[~] Palette color-index map coding now shares the exact current-libaom neighbor weights, stable color ordering, five context classes, first-index uniform code, and diagonal wavefront between encoder costing, encoder writing, and decoder parsing. The decoder's stack-allocated context scores are explicitly cleared before accumulation, removing an invalid dependency on uninitialized stack contents. Costing and writing use a closed generic operation while the shared driver owns traversal and context derivation, so the semantic operations remain independent of map layout and tail handling. The path adds no retained state or per-call managed allocation. Its allocation regression now runs one complete unmeasured hot-path window before measuring an independent 1,000-call steady-state window, so tiered-runtime transitions cannot make the full parallel suite report a one-time allocation as a recurring operation cost. Twelve focused map, exact palette decode, padding, trailing-bit, and allocation cases pass; all 1,941 entropy cases and all 8,991 HEIF/AV1 cases pass direct net11 Release VSTest. The exact Release rebuild remains at 1,005 warnings and zero errors. Production payloads remain unchanged because palette selection is still disabled; retained colors, neighbor caches, index-map storage, candidate generation, and production palette mode decision remain incomplete.
-
[~] Retained encoder palette state and production palette writing now mirror current libaom's 50-byte palette-mode contents, separate luma and shared-chroma sizes, three eight-color planes, above-and-left sorted cache, 64-sample above-cache boundary, mode contexts, palette colors, color-index maps, and syntax order. The current block keeps one inline value in the reusable superblock workspace; only the 4x4-granularity top and left picture edges retain copies for later blocks. For a 3840x2160 tile these edges occupy about 73.4 KiB instead of about 6.2 MiB for a 50-byte palette value attached to every 8x8 mode allocation. Luma and chroma index maps share one lazily allocated 32 KiB owner containing two 128x128 maps, so the palette-disabled production path retains no map owner. The compact final-block workspace falls from about 10.3 KiB to about 8.3 KiB. The writer caps map traversal to the coded plane count, writes maps before transform syntax, and publishes palette edges only after the current block has consumed preceding contexts. Eight focused size, alignment, ownership, cache-boundary, round-trip, map-consumption, and edge-publication cases pass; all 114 palette cases, all 1,942 entropy cases, and all 8,996 HEIF/AV1 cases pass direct foreground net11 Release VSTest. The exact net11 Release rebuild reports 1,050 solution warnings and zero errors; Roslynk reports zero compiler errors and no diagnostics in the touched files. The current-main reference is
a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727. Production payloads remain unchanged because palette candidate generation is still disabled; that live rate-distortion search is the next checkpoint. -
[~] Luma palette clustering now follows current libaom's one-dimensional search primitive exactly: equal-interval midpoint initialization, first-color tie order, rounded centroid means, deterministic empty-cluster replacement, the 50-iteration limit, and retention of the preceding state when distortion increases. Nearest-color assignment improves on libaom's AVX2 implementation by dispatching Vector512, Vector256, Vector128, then scalar through ImageSharp's shared vector-count helpers. The primitive uses only bounded stack scratch and introduces no allocator rent, managed array, or per-row copy. Three independent tests cover exact centroid convergence, initialization order, 12-bit nearest-color distortion, destination bounds, and every hardware-intrinsic tier. The complete AVIF set passes 8,930 of 8,930 cases and the HEIF set passes 230 of 230 cases through direct foreground net11 Release VSTest. The exact net11 Release rebuild reports 1,050 solution warnings and zero errors, and Roslynk reports zero compiler errors. Candidate enumeration, palette-cache snapping, transform RD selection, and production activation remain in the open luma-palette checkpoint.
-
[~] Live luma palette selection now follows current libaom's dominant-color and one-dimensional K-means candidate families, cache-bias threshold, sorted duplicate removal, active-edge map extension, and strict winner tie order. It improves on speed-configured libaom by evaluating both candidate families at every legal 2-through-8 size without early pruning, then exhaustively evaluates every legal transform using the existing SIMD prediction, residual, transform, quantization, and reconstruction operators. Candidate storage remains bounded stack memory; the reusable 32 KiB map owner is allocated only when an eligible block enters palette search. A production tile test proves that full 8x8 and clipped 5x3 blocks at 8 and 12 bits select exact colors and indices, extend the visible edges through coded padding, reconstruct every sample without coefficients, and emit a nonempty tile. The complete 57-case intra-superblock set, 8,931-case AVIF set, and 230-case HEIF set pass direct foreground net11 Release VSTest. The exact Release test-project build reports 1,992 baseline warnings and zero errors; Roslynk reports zero compiler errors and no touched-file analyzer warnings. Production frame activation remains gated until chroma palette mode and its rate accounting are complete.
-
[~] Paired chroma palette clustering now preserves current libaom's squared two-component distance, first-centroid tie order, independently rounded U/V means, paired deterministic empty-cluster replacement, preceding-state retention on increased distortion, and 50-iteration limit. Keeping the source planes separate avoids interleave/deinterleave copies and improves on libaom's AVX2 ceiling with Vector512, Vector256, Vector128, then scalar dispatch through ImageSharp's shared vector-count helpers. Three independent tests cover exact paired convergence, midpoint initialization, 12-bit distance and index parity, untouched destination bounds, and every intrinsic tier. The exact Release test-project build reports 1,992 baseline warnings and zero errors; the focused three-case set, complete 8,934-case AVIF set, and complete 230-case HEIF set pass direct foreground net11 Release VSTest. Roslynk reports zero compiler errors and no touched-file analyzer warnings. Candidate integration and production activation remain in the open chroma-palette checkpoint.
-
[~] Live paired chroma palette selection now follows current libaom's complete 2-through-8 color-size search, U-plane neighbor-cache snapping, stable U-ordered color pairs, shared U/V index map, implicit DCT-DCT transform, and strict rate-distortion winner replacement. It improves on speed-configured libaom by applying no early header-cost pruning, keeps planar U/V source data separate, and reuses the SIMD-first prediction, residual, transform, quantization, and reconstruction operators without allocator-backed candidate storage. The production tile regression proves both palette-mode probability branches, exact paired colors and indices, coefficient-free reconstruction, and nonempty syntax. The complete 58-case intra-superblock set, 8,935-case AVIF set, and 230-case HEIF set pass direct foreground net11 Release VSTest. The exact Release test-project build reports 1,992 baseline warnings and zero errors; Roslynk reports zero compiler errors and no touched-file analyzer warnings. Production frame activation remains the next checkpoint.
-
Production screen-content activation now matches current libaom's default good-quality detector: it scans only complete 16x16 luma blocks, normalizes palette samples to eight bits, admits 2-through-4-color blocks, and uses the reference's strict greater-than-ten-percent frame-area threshold. The same pass accumulates centered sums and squared sums at native precision, applies libaom's exact 10-bit and 12-bit variance rounding, and enables intra-block copy only when positive rounded per-pixel variance exceeds its strict one-twelfth frame-area threshold. A 256-bit stack bitset and fifth-color early exit replace libaom's larger per-block histogram without a second source scan or allocation. The adaptive sequence flag remains enabled and both frame flags are fixed before picture-state allocation. Focused regressions prove strict palette-threshold equality, high-bit-depth normalization, the exact variance rounding boundary, five-color rejection, emitted frame-header activation, actual production IBC selection, and production decode. The exact Release test-project build reports 1,992 baseline warnings and zero errors; all 9,242 non-HEVC HEIF/AV1 cases pass direct foreground net11 Release VSTest. Current-main
aomdecata40ed1ea9e4ecc3df58a5bccb76623f2c94ae727accepts all 31 payloads regenerated by the current test tree, including an actual IBC-coded 328x16 stream with decoded MD5677435e5af39c930af1178f91c34af6a. Roslynk reports zero compiler errors and no touched-file analyzer warnings. -
[~] Intra-block-copy rate accounting now uses the live frame-local flag and displacement-vector distributions without copying or adapting either context during candidate measurement. Displacement-vector costing and writing share one closed symbol operation over the exact current-libaom joint, sign, magnitude-class, class-zero, and integer-offset syntax; final mode evaluation applies libaom's 120/128 displacement-rate weight with nearest-integer rounding. Independent fixed costs cover all four joint states, both signs, class zero, and large offset classes before adaptive writes, followed by an encoder/decoder round trip through the same sequence. Encoder and decoder reference-vector derivation now share the exact eight-candidate spatial scan, independent nearest and outer-region ranking, top-right partition geometry, clamping, and tile-relative fallback. Selected vectors use a naturally aligned pair of signed 16-bit components packed into the existing picture-state owner only when intra-block copy is permitted; a 3840x2160 frame retains 130,560 vectors in 510 KiB while leaving the compact 8-byte mode allocation unchanged. The tile writer derives the same reference and emits the retained vector without another allocation or copy. Coefficient costing and writing now select the inter transform sets and frame-local probability tables required by intra-block copy; independent tests verify every legal symbol against the exact default inter distribution and round-trip full and reduced sets from 4x4 through 32x32. Legal 8x8 hash discovery now indexes every visible source origin, including unaligned origins, in libaom's coarse-to-fine insertion order with the same 256-candidate bucket cap. A separable rolling hash fills one packed picture-lifetime workspace before reconstruction, then reuses that workspace for integer candidate links; exact wide or SIMD block comparison rejects hash collisions, and SIMD variance uses libaom's eight-bit normalization at 8, 10, and 12 bits. Power-of-two bucket arrays scale down with small images and stop at the reference's 16-bit limit, avoiding libaom's fixed six-size pointer table; the 3840x2160 search index occupies about 32.2 MiB and introduces no additional owner or frame copy. Above and left search rectangles, integer displacement legality, strict tie order, and live raw displacement rate follow current libaom. Motion-candidate ranking uses libaom's undiscounted probability cost and exact variance-domain error-per-bit scaling, separately from the later 120/128 final-mode discount. The allocation-free full-pixel core now follows current libaom's NSTEP search: it clamps the spatial reference to each legal region, traverses the fixed 15-stage radii and site order, skips equivalent centered 210-pixel stages, repeats progressively shorter paths, and compares their winners in the normalized variance domain. Paths above the speed-zero screen-content threshold continue through libaom's 256-pixel, one-pixel-step exhaustive mesh. Four adjacent byte or high-bit-depth candidates share each SIMD source load, strict row-major tie ordering is retained, and the final legal tail column remains searchable where libaom's current four-wide remainder loop omits it. Byte and high-bit-depth operators compute each 8x8 absolute difference with Vector128 before scalar fallback; high-bit-depth SAD remains in its native sample scale while its quantizer-derived rate multiplier uses libaom's normalized AC step. Production mode decision now derives the same spatial displacement reference used by the writer, deduplicates hash and full-pixel finalists in search order, and evaluates every surviving vector through complete luma and chroma transform RD. This intentionally improves on libaom's preliminary-error pruning by permitting a hash and pixel finalist from the same search region to compete using their final syntax and reconstruction costs. Prediction is prepared once per plane and vector, including integer or half-sample chroma phase, then reused across every legal inter transform without an allocator rent or frame copy. The joint comparison includes the live intra-block-copy flag, discounted displacement rate, skip flag, coefficient syntax, and normalized Y/U/V distortion; an empty transform alternative can win only when its complete skip cost is strictly lower, while conventional intra and earlier vectors retain tie precedence. Winning reconstruction, coefficients, transform state, DC modes, cleared palette/filter/CfL state, and displacement are copied once into the existing retained stores. Production regressions force the path at 8 and 12 bits and force 4:2:0 horizontal half-sample chroma with an unaligned reference. The former bulk local workspace occupied 2.75 KiB for byte samples or 3.375 KiB for high-bit-depth samples. Prediction, candidate, winning reconstruction, residual, and coefficient scratch now occupy one naturally aligned 3.125 KiB extension of the existing frame-reused block-workspace owner, matching libaom's reusable macroblock-scratch lifetime without adding an allocation; only the 128-byte reference, weight, and finalist arrays remain on the stack. The net11 Release solution build reports zero errors; all 2,082 focused transform, entropy, intra-block-copy, intra-superblock, and frame-encoder cases and all 9,242 non-HEVC HEIF/AV1 cases pass through direct foreground VSTest, with tiered compilation disabled only for the full allocation-sensitive suite. Adaptive production activation is complete, and the emitted frame flag remains authoritative for the complete frame rather than being invalidated after tile coding.
-
The expanded checkpoint exposed a pre-existing transform-block test that asserted uninitialized pooled padding was zero. The test now initializes the complete physical luma plane with a sentinel and proves the block operation leaves both adjacent padding samples unchanged. The exact net11 Release rebuild remains at 1,005 baseline warnings and zero errors, the focused allocator-order set passes 30 of 30 cases, and the complete HEIF/AV1 namespace passes 8,859 of 8,859 direct VSTest cases with zero failures or skips.
-
Combined-frame OBU output now counts the byte-aligned frame and tile-group headers, non-final tile-size fields, and owned tile payloads before emitting the OBU size. It retains only the small allocator-owned header scratch and writes each entropy-coded tile span directly from its detached owner, removing the second file-sized allocator rent and complete-payload copy. A 64 KiB regression proves exactly one sub-payload-sized byte rent with a balanced return and verifies the exact streamed tile tail; the existing two-tile round trip proves size-prefix and ordering parity. The focused writer and production-frame set passes 32 of 32 direct net11 VSTest cases, current-main
aomdecaccepts all 29 generated native-format payloads, and the complete HEIF/AV1 namespace passes 8,860 of 8,860 cases with zero failures or skips. -
Finalized fixed-block decisions now set the block-level transform-skip flag only when every retained luma and coded chroma transform has zero EOB, matching current libaom's conjunction of per-plane skip state. The previous always-false flag produced legal but redundant non-skip and zero-coefficient syntax. Monochrome and 4:2:0 regressions prove both branches from actual coefficient state; the focused decision and production-frame set passes 32 of 32 direct net11 VSTest cases. Current-main
aomdecaccepts all 29 regenerated payloads, the recorded decoded-frame MD5s are unchanged, and affected 16x16 constant 8-bit and 10-bit payloads are one byte smaller. The complete HEIF/AV1 namespace passes 8,862 of 8,862 cases with zero failures or skips. -
Operation-wide allocation tracking now exercises a real 64x64 12-bit 4:4:4 frame through packed-pixel conversion, both native frame owners, picture and coefficient state, reusable block workspaces, entropy coding, OBU framing, and a non-seekable destination. It proves exactly one 60 KiB tile-output reservation from current libaom's all-intra 2.5x rule and balanced exactly-once returns for every tracked allocation before the operation completes. The focused ownership case passes 1 of 1 and the complete HEIF/AV1 namespace passes 8,863 of 8,863 direct net11 VSTest cases with zero failures or skips.
-
HEIF box offsets are now counted from the start of the encoded file instead of reading
Stream.Position. This preserves ISO BMFF file-relativeilocoffsets when the destination begins at a nonzero position and permits non-seekable output. Decoder item extents and image-sequence chunk offsets now resolve from that same file origin rather than the backing stream origin. Real legacy-JPEG HEIF round trips cover non-seekable output and a prefixed destination, while current-position AV1 decode covers both a still item and a five-frame sequence. All 96 encoder/decoder cases and all 38 sequence-parser cases pass direct net11 Release VSTest; the Release build remains at the established 1,005-warning baseline with zero errors. -
[~] Effort-six uniform luma transform selection now compares the winning ordinary spatial mode as one 8x8 transform against four raster-ordered 4x4 transforms. Each 4x4 transform searches every legal type with coefficient contexts derived from the already retained transform edges and the preceding trial blocks, while reconstructed top-right and bottom-left references follow production coding order. The strided transform operator writes each candidate directly into its 8x8 reconstruction mosaic. Prepared prediction and residual data are reused across transform trials, and the existing aligned block-workspace owner retains four final transform states, local coefficient contexts, coefficients, and compact trial reconstruction; no allocator rent, managed array, best-candidate re-transform, or full-block intermediate copy was added. Non-skipped intra-block copy under
TX_MODE_SELECTemits and costs the unsplit variable-transform root required by current libaom, while skipped intra-block copy emits no transform-partition symbol. Uniform transform-size contexts use coding-block extents for intra-block-copy neighbors and residual-transform extents for intra neighbors. The focused contract passes 21 of 21 direct net11 Release VSTest cases, including an independently decoded stream that proves at least one real four-transform luma block. The complete non-HEVC HEIF/AV1 namespace passes 9,294 of 9,294 cases with zero failures or skips. The exact Release build remains at 1,005 baseline warnings and zero errors, Roslynk reports no touched-file diagnostics, and current-mainaomdecaccepts both generated effort-six streams with decoded-frame MD5 values2dd1cbe449fe2d0471dc2c15c50acb69and677435e5af39c930af1178f91c34af6a. Transform-size integration for filter-intra and palette, broader joint mode/transform refinement, partition search, and effort-dependent pruning remain.
7. Write complete AVIF output
- [~] The encoder-side AV1 codec configuration is now derived directly from the encoded sequence header and writes the fixed four-byte
av1Crecord with emptyconfigOBUs. The image payload retains the required sequence header, so the property introduces no sequence-header allocation, retention, or copy. Four production-header cases cover main, high, and professional profiles; 8-, 10-, and 12-bit precision; monochrome, 4:2:0, 4:2:2, and 4:4:4 sampling; exact fixed bytes; decoder reparsing; and header/property equivalence through direct net11 Release VSTest. Property-container emission and public AVIF activation remain open. - [~] AV1 image properties now write
ispe,pixi,av1C,colr, andauxCin current AVIF item order. Onlyav1Cis essential; color and alpha items retain independent property sets and the registered alpha auxiliary type. The property container reacquires its span after nested expansion before patchingipco, removing the prior stale-buffer write, and selects compact or 15-bitipmaindices from the property count rather than the unrelated item count. A forced-growth color-plus-alpha case validates every property payload and association byte; a separate 43-item, 129-property case proves indices 127 through 129 and the extended essential bit. Both pass direct foreground net11 Release VSTest. Complete AVIF assembly remains open. - [~] Explicit public AV1 encoding now writes a still-image AVIF with
avifmajor brand, compatibleavif,mif1, andmiafbrands, one primary color item, an optional alpha auxiliary item,auxlfrom alpha to color, independent item properties, absolute version-oneilocextents, and a sharedmdat. Quality uses current libaom's quantizer-to-qindex mapping with public quality 100 deliberately clamped from lossless qindex 0 to qindex 4. Effort controls the implemented search stages, and the resolved value is required explicitly by every internal frame, tile, and mode-decision operation rather than repeated as optional defaults. Encoder options take precedence over source metadata for 8-, 10-, and 12-bit monochrome, 4:2:0, 4:2:2, and 4:4:4 output. Alpha derives from the source pixel type without scanning pixels, and incompatible identity-matrix metadata is normalized without mutating the source image. - [~] The production path writes color and alpha payloads sequentially through allocator-backed chunked storage, supports non-seekable and prefixed destinations, and does not materialize a complete file or payload copy. Uniform encoder-side
pixidepth is written directly without allocating per-item channel-depth arrays; decoder-side non-uniform channel depths remain supported. The Release test project builds with zero errors, all 39 encoder cases pass, the complete non-HEVC HEIF namespace passes 9,277 of 9,277, and current official libaom accepts all 47 generated payloads. - Still-image AVIF metadata preservation now writes an unrestricted ICC
colr/profproperty before the independentcolr/nclxproperty, Exif and XMP as separatemdatitems, and onecdscrelationship from each metadata item to the primary color item. Exif stores the exact big-endian TIFF-header offset required by the HEIF item syntax; XMP uses themimeitem type andapplication/rdf+xmlcontent type. Existing ICC and XMP storage is read synchronously and copied once into final encoder storage rather than cloned into an intermediate array.SkipMetadatasuppresses all three profile types while retaining the CICP values required to describe the encoded planes. The same option now reaches legacy JPEG payloads, whose encoder no longer writes application profiles or comments when metadata is disabled. - Exact container tests verify every emitted item declaration, name, MIME content type,
cdscrelationship, Exif offset and payload, XMP payload, ICC/CICP property order, compact association byte, propertyless metadata exclusion, decoded profile value, and bothSkipMetadatabranches. The final HEIF encoder set passes 44 of 44 and the complete JPEG encoder set passes 257 of 257 through direct foreground net11 Release VSTest. The complete non-HEVC HEIF namespace passes 9,282 of 9,282 with no failure, crash, or detached test host, and current official libaom accepts all 47 current generated AV1 payloads. - A code-wide production HEIF/AV1 stack-storage audit, excluding HEVC, removed every block-sized, variable-length, or repeatedly nested scratch buffer. Spatial luma and chroma, filter-intra, chroma-from-luma, luma and chroma palette selection, and K-means iteration now use typed views over 642 signed-integer elements, about 2.51 KiB, at the start of the existing 3.125 KiB IBC region. Those searches are sequential for one block, so the block-workspace owner does not grow and no rent, copy, or additional lifetime is introduced. CDEF directions, variances, and its 64-entry block list now append 1 KiB to the existing bounded operation owner instead of occupying hidden inline or explicit stack arrays. No remaining
stackallocdepends on block dimensions, sample count, or runtime length; the largest remaining individual span is 128 bytes, and the remaining sites are fixed syntax, SIMD-lane, filter-tap, plane-metadata, or small candidate storage. The exact-owner test now proves the mode, palette, and IBC views share one allocation. Roslynk reports zero compiler errors and no diagnostics in the changed files, the Release test-project build completes with the established 1,992 warnings and zero errors, 81 of 81 focused cases pass, and the complete non-HEVC HEIF/AV1 namespace passes 9,282 of 9,282 through one foreground net11 VSTest run. - [~] Write the correct AVIF file type, item information, locations, references, properties, AV1 configuration, dimensions, color, alpha, metadata, and media data.
- [~] Support single images, alpha auxiliary images, grids, multiple extents, and bounded image sequences in the final public scope.
- Preserve ICC, Exif, and XMP according to encoder options.
- [~] Write CICP, range, chroma position, bit depth, and subsampling values that match the encoded planes.
- Apply orientation and clean-aperture behavior consistently with ImageSharp encoder conventions.
- [~] Stream output through allocator-backed chunked storage without file-sized copies or ToArray materialization.
Encoder exit gate:
- Current-main libaom accepts every produced AV1 payload.
- Lossless output is exact at native-plane and final-pixel precision.
- Lossy output demonstrates recorded quality and effort tradeoffs with absolute size, quality, timing, and allocation evidence.
- 8, 10, and 12-bit monochrome, 4:2:0, 4:2:2, and 4:4:4 outputs pass.
- Alpha, grids, metadata, color profiles, transforms, and bounded sequences pass.
- ImageSharp decode of its own output is supplemental coverage only, never the sole oracle.
- Public encoding no longer throws for a supported AV1 request.
- Focused Release and FeatureTestRunner verification passes with exact recorded evidence.
Architecture rules
- Follow the JPEG color-converter operator architecture exactly.
- Each distinct prediction traversal owns a family-named predictor type.
- The family .Operator.cs file defines the nested static operator contract.
- Each semantic readonly struct belongs to that owner and implements concrete scalar, Vector128, Vector256, and Vector512 arithmetic for the shared traversal.
- Do not place a distinct predictor beneath a broad Av1IntraPredictor or Av1InterPredictor.
- Do not create semantic forwarding wrappers, top-level operator types, hardware-width-named operator types, CRTP contracts, or one file containing unrelated semantic operators.
- Forward transforms belong to Av1ForwardTransformer and its semantic operator files.
- Inverse axis transforms belong to Av1Inverse2dTransformer and its semantic operator files.
- Reconstruction output operators belong to Av1InverseTransformer.
- Shared lane primitives belong only in explicitly named Operations types.
- Dispatch from widest to narrowest supported SIMD width, then execute one scalar tail.
- Keep codec execution sequential. Do not add parallel execution inside the codec.
- Do not allocate per row, block, transform, scanline, or SIMD tail.
- Use ImageSharp allocators and pools. Do not use ToArray to cross an ownership boundary.
- On internal types, use public members when other types consume them; reserve private members for type-local behavior.
- Use established ImageSharp test data, allocator tracking, FeatureTestRunner, and reference-image comparison APIs. Do not build custom substitutes.
- Public XML documentation describes observable behavior only.
- Inline comments explain the current-libaom numerical rule, ownership boundary, edge extension, entropy ordering, or SIMD shape at technically complex points.
- Every multiline statement or declaration is followed by vertical whitespace.
- Do not edit .gitattributes directly.
- Do not install or download tools without explicit permission.
Final verification matrix
- Release source build: net10.0, zero errors.
- Release source build: net11.0, zero errors.
- Scoped semantic inspection: zero compiler errors attributable to this work.
- Focused decoder syntax, reconstruction, ownership, presentation, and malformed-input tests.
- Focused encoder syntax, payload, container, precision, ownership, and option tests.
- FeatureTestRunner coverage for normal, narrower SIMD tiers, and scalar fallback.
- Constrained multi-group allocator coverage with balanced exactly-once returns.
- Exact native-plane comparisons against current-main libaom.
- Established final-presentation comparisons at the target pixel precision.
- Scoped StyleCop and vertical-whitespace inspection.
- No stale unsupported capability claims or removed-code references.
- No restore-source failures, background test hosts, detached processes, or crash-report popups.
- .gitattributes unchanged.
- git diff --check clean.
- Documentation records exact commands, counts, current-main reference revision evidence, and results.
- Commit only after the relevant checkpoint is genuinely complete.
- Do not push.