From 4c7880dbe23d390718ad7bff417c28388b2cf15a Mon Sep 17 00:00:00 2001 From: James Jackson-South Date: Sat, 5 Sep 2026 19:08:20 +1000 Subject: [PATCH] Record AV1 retained-state and film-grain source audit --- HEIF_IMPLEMENTATION_PLAN.md | 50 +++++++++++++++++++ .../Pipeline/FilmGrain/Av1FilmGrainNoise.cs | 17 +++---- .../Pipeline/FilmGrain/Av1FilmGrainOverlap.cs | 10 ++-- 3 files changed, 63 insertions(+), 14 deletions(-) diff --git a/HEIF_IMPLEMENTATION_PLAN.md b/HEIF_IMPLEMENTATION_PLAN.md index 322d99c6b4..3084fdbf43 100644 --- a/HEIF_IMPLEMENTATION_PLAN.md +++ b/HEIF_IMPLEMENTATION_PLAN.md @@ -320,6 +320,56 @@ Frame/block RD and decoder filter follow-up after `aa2ecf690`: retains bottom lines for its worker-capable traversal. This inspection establishes no decoder-wide or SIMD completeness claim. Decoder CDEF storage remains an operation-scoped owner, not native reusable worker state. +Retained-state and cost-policy follow-up after `ef8b1a823`: + +- `Av1EncoderTransformBlockState.cs:12-43` uses four bytes for EOB and byte-sized transform type, leaving + one padding byte. Reference `av1/encoder/encodetxb.c:593-625,714-730` records the neighboring skip context + in bits 0-3 and DC-sign context in bits 4-5 beside each EOB. Its packer consumes those retained contexts + (`encodetxb.c:296-306,410-421`). The managed structure can represent that state without increasing its size, + but storing it alone would not implement deferred packing; no unused state was added. +- Reference palette-token allocation is conditional on non-statistics coding with screen-content tools allowed + (`av1/encoder/encodeframe.c:1413-1430`). `tokenize.h:105-135` reserves up to two full-resolution planes in + maximum-superblock-rounded storage. Tokens retain the selected context and color-order rank, including the + first raw index (`tokenize.c:174-225,264-278`, `bitstream.c:353-368`); keeping only palette colors is insufficient. +- `Av1EncoderPictureBuffer.Reset` clears the complete mode grid and packed state (`:344-363`), so it cannot + be reused as the boundary between analysis and packing. Selected prediction fields remain in the reusable + `Av1EncoderBlockStruct` workspace, while `Av1EncoderBlockModeInfo` retains only its smaller neighbor subset. + These lifetimes must be reconciled together with palette tokens, selected MV state, and coefficient contexts. +- Native cost defaults are explicit controls as well as speed features: `av1/av1_cx_iface.c:391-394,550-553` + differs between default and realtime configurations. `rd.c:724-758,824-851` combines controls with speed policy + and initializes frame costs; `encodeframe_utils.c:1629-1689` suppresses block refresh when CDF updates are + disabled. No new managed effort mapping or isolated cost-refresh threshold was introduced. + +Film-grain decoder source comparison after `ef8b1a823`: + +- The complete template generation, random state, autoregression, scaling interpolation, overlap traversal, + noise application, and native-sample load/store paths were compared with `av1/decoder/grain_synthesis.c`. + Managed `Av1FilmGrainDecoder.cs:874-1172,1204-1231` matches the represented rules in native `:429-629`; + all 2,048 Gaussian entries also match exactly. No new arithmetic defect was established in this comparison. +- Managed noise application processes chroma before luma (`Av1FilmGrainNoise.cs:109-174`), retaining ungrained + luma for chroma scaling. Native `grain_synthesis.c:685-745,803-862` uses that same ordering. The two-component + luma average is horizontal only; vertical chroma subsampling selects a row rather than averaging two rows. + Restricted identity-matrix chroma uses luma's upper endpoint, and high-depth lookup interpolates below entry 255. +- `Av1FilmGrainDecoder.cs:698-857` and native `grain_synthesis.c:1252-1376` exclude already-applied overlap + strips and retain right/bottom grain boundaries. The managed plane span retains existing padded storage + (`:208-216`), including addresses for empty interiors at clipped edges; no new guard or scratch plane was added. +- Grain presentation preserves references in `Av1Decoder.cs:891-909,945-1002`: refreshed frames receive a + separate presentation copy, and unreferenced shown frames can be grained in place. `CopyVisibleTo` also copies + active geometry (`Av1FrameBuffer.cs:248-258`). This source trace does not complete the decoder-wide lifetime audit. +- SIMD dispatch remains an open architecture/performance issue. `Av1FilmGrainNoise.cs:180-238,388-522,947-955` + uses AVX2 gather or a high-depth-only portable path with separate width overloads. `Av1FilmGrainOverlap.cs:142-175` + additionally gates 512-bit processing on `Vector.Count`. Neither gate has fresh end-to-end evidence here. + Comments claiming that scalar reads are slower, interpolation repays them, or `Vector` establishes processor + execution width were corrected. Runtime dispatch and arithmetic were not changed; no improvement is claimed. +- After the final comment edit, the Release .NET 11 build completed with zero errors and 1,009 existing warnings; + Roslynk reported zero compiler errors. Serialized Visual Studio VSTest passed 5/5 focused film-grain/reference + cases in 8.4512 seconds (`film-grain-audit.trx`), including existing hardware-fallback and constrained-allocation checks. +- Current optimized libaom regenerated seven still references and the two ten-frame official sequence references. + All 3,113,847 decoded samples match the retained references exactly: maximum error 0, zero samples exceeding one. + Per-frame/per-plane results are in temporary `film-grain-comparison.json`. Those references are also used by the + passing managed tests. This is bounded same-bitstream decoder evidence, not complete conformance or encoder parity. + No benchmark ran, and no fixture, native integration, or generated comparison output was added to the repository. + Range-writer output-capacity correction, verified after `93aba785f`: - `Av1SymbolWriter.cs:213-214,339` before correction sliced a fixed initial allocation for finalization diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/FilmGrain/Av1FilmGrainNoise.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/FilmGrain/Av1FilmGrainNoise.cs index 24f5dd87ed..6293db935e 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/FilmGrain/Av1FilmGrainNoise.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/FilmGrain/Av1FilmGrainNoise.cs @@ -15,9 +15,8 @@ namespace SixLabors.ImageSharp.Formats.Heif.Av1.Pipeline.FilmGrain; /// /// SIMD lanes follow consecutive samples within one plane row. Grain values and native samples are widened to signed /// 32-bit lanes before the scaling-table lookup and fixed-point addition, then clipped and narrowed only once. AVX2 -/// uses indexed gathers for the 256-entry scaling table. Portable 128-bit traversal is retained where high-bit-depth -/// interpolation provides enough arithmetic to offset its scalar table reads; all remaining columns use the identical -/// scalar equation. +/// uses indexed gathers for the 256-entry scaling table. The current portable 128-bit dispatch handles high-bit-depth +/// interpolation with scalar table reads; all remaining columns use the same scalar equation. /// internal static class Av1FilmGrainNoise { @@ -417,8 +416,8 @@ internal static class Av1FilmGrainNoise int maximum) where TSample : unmanaged { - // Chroma uses the same indexed scaling-table constraint as luma, so gather support determines the primary - // width and the portable path remains restricted to workloads that amortize scalar table reads. + // Chroma shares luma's dispatch: AVX2 gathers scaling values, while the portable vector path is currently + // enabled only for high-bit-depth interpolation. This policy does not establish which path is faster. if (Avx2.IsSupported) { ApplyChroma( @@ -941,16 +940,16 @@ internal static class Av1FilmGrainNoise } /// - /// Determines whether portable vector arithmetic repays the cost of scalar scaling-table reads. + /// Determines whether the current dispatch enables the portable vector traversal. /// /// The decoded sample bit depth. /// Whether to use the portable vector traversal. [MethodImpl(MethodImplOptions.AggressiveInlining)] private static bool CanVectorizeWithoutGather(int bitDepth) { - // Portable Vector128 has no indexed table load. At eight bits, assembling each scaling vector from four - // scalar reads is slower than the complete scalar loop; high-depth interpolation contains enough arithmetic - // to amortize those reads. The AVX2 path uses native gather and remains the primary traversal at every depth. + // This path assembles four scale values from scalar lookups. High-depth samples additionally interpolate + // between table entries. The current depth gate requires end-to-end evidence before it can be justified + // as a performance policy; it does not express a numerical requirement of film-grain synthesis. return bitDepth > 8 && Vector128.IsHardwareAccelerated; } diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/FilmGrain/Av1FilmGrainOverlap.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/FilmGrain/Av1FilmGrainOverlap.cs index 3cd2be0b66..a23bc17227 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/FilmGrain/Av1FilmGrainOverlap.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/FilmGrain/Av1FilmGrainOverlap.cs @@ -13,7 +13,7 @@ namespace SixLabors.ImageSharp.Formats.Heif.Av1.Pipeline.FilmGrain; /// /// /// Vertical boundaries contain only one or two strided columns and use the fixed scalar kernels. Horizontal boundaries -/// are contiguous and progress from the runtime's preferred native width through smaller vector widths before the +/// are contiguous and progress through the enabled vector widths before the /// scalar tail. Every lane applies the same Q5 overlap weights, rounding offset, and signed grain clamp. /// internal static class Av1FilmGrainOverlap @@ -43,8 +43,8 @@ internal static class Av1FilmGrainOverlap int minimum, int maximum) { - // Each row contributes only one or two strided samples. Gather plus scalar scatter would do more work than - // the fixed scalar kernel, while the horizontally contiguous boundary below benefits directly from SIMD. + // Each row contributes one or two strided samples. This traversal handles those columns directly; + // the horizontal traversal below groups contiguous samples into vector lanes. if (width == 1) { for (int row = 0; row < height; row++) @@ -151,8 +151,8 @@ internal static class Av1FilmGrainOverlap { int column = 0; - // Vector exposes the runtime's preferred native width. This avoids selecting split 512-bit operations on - // machines whose execution resources are 256 bits wide while retaining a native 512-bit traversal elsewhere. + // The current 512-bit gate also requires Vector to expose sixteen int lanes. That is a dispatch choice, + // not evidence of the processor's execution width or of faster overlap processing. if (Vector512.IsHardwareAccelerated && Vector.Count == Vector512.Count) { column = Blend(left, right, destination, width, column, leftWeight, rightWeight, minimum, maximum, Vector512.Zero);