From bd66e473827d2b1b8000b0d73a2292a8fe0a421f Mon Sep 17 00:00:00 2001 From: James Jackson-South Date: Thu, 3 Sep 2026 05:29:45 +1000 Subject: [PATCH] Activate AV1 palette coding adaptively --- HEIF_IMPLEMENTATION_PLAN.md | 3 +- .../Heif/Av1/Pipeline/Av1FrameEncoder.cs | 4 + .../Av1/Pipeline/Av1ScreenContentDetector.cs | 122 ++++++++++++++++++ .../Formats/Heif/Av1/Av1EncoderFrameTests.cs | 122 ++++++++++++++++++ 4 files changed, 250 insertions(+), 1 deletion(-) create mode 100644 src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1ScreenContentDetector.cs diff --git a/HEIF_IMPLEMENTATION_PLAN.md b/HEIF_IMPLEMENTATION_PLAN.md index 3aae2daf4a..5548e80b51 100644 --- a/HEIF_IMPLEMENTATION_PLAN.md +++ b/HEIF_IMPLEMENTATION_PLAN.md @@ -825,7 +825,7 @@ Encoder verification contract: - [~] A non-owning encoder-frame view now separates visible conversion regions from coded regions and performs complete left, top, right, bottom, and corner extension across each bordered plane. Current libaom uses 8-sample-aligned coded dimensions, a 32-sample-aligned luma stride with chroma stride derived from it, and a 64-pixel luma border for non-resized all-intra encoding. One operation-ready frame owner now rents the aligned Y, U, and V storage contiguously, exposes non-owning `Buffer2D` plane views, and returns the rent exactly once. A 4K 4:2:0 frame occupies about 13.0 MiB at 8-bit or 26.0 MiB at 10/12-bit; source and reconstruction therefore remain distinct frame owners rather than adding a full-frame copy. The corrected tests use this real ownership path and verify the exact 54 KiB 64x64 4:2:0 rent. The frame-encoder operation now instantiates matching source and reconstruction owners with ordinary `using` lifetimes and converts packed pixels directly into the source owner before extension. - [~] Temporal delimiter, sequence header, frame header, combined-frame tile-group writing, and an internal reduced-still-picture frame operation now exist locally. The remaining required metadata, padding, multi-tile, option, and public encoder paths are not complete. - [~] Implement superblock and partition analysis for every permitted block size and partition. The current baseline deliberately splits every in-frame node to 8x8 blocks and records decisions in current-libaom writer preorder; block-size selection and non-split partition analysis remain. -- [~] Implement intra mode search, palette, filter intra, chroma-from-luma, and intra-block copy decisions. Live luma search now covers all 13 zero-angle base modes and all six nonzero adjustments for each of the eight directional modes. Joint spatial chroma search covers the same 61 candidates, combines both chroma planes in one rate-distortion decision, and preserves the winning shared angle adjustment. Chroma-from-luma now searches the complete signed alpha alphabet from reconstructed luma and retains its joint U/V syntax. Filter-intra now searches all five predictors after ordinary luma modes. Palette entropy, retained state, production syntax, and exhaustive luma and paired chroma palette selection are complete; intra-block-copy mode decisions remain. +- [~] Implement intra mode search, palette, filter intra, chroma-from-luma, and intra-block copy decisions. Live luma search now covers all 13 zero-angle base modes and all six nonzero adjustments for each of the eight directional modes. Joint spatial chroma search covers the same 61 candidates, combines both chroma planes in one rate-distortion decision, and preserves the winning shared angle adjustment. Chroma-from-luma now searches the complete signed alpha alphabet from reconstructed luma and retains its joint U/V syntax. Filter-intra now searches all five predictors after ordinary luma modes. Palette entropy, retained state, production syntax, exhaustive luma and paired chroma palette selection, and adaptive production activation are complete; intra-block-copy mode decisions remain. - [ ] Implement inter mode search for bounded sequences, including reference selection and the decoder-supported inter tools. - [~] Current-libaom `av1_quantize_fp_no_qmatrix` arithmetic is implemented as a closed generic forward-quantizer family with Vector512, Vector256, Vector128, and scalar paths, raster-order output, coded 64-point coefficient limits, and scan-order EOB selection. Transform search, coefficient optimization, and lossless behavior remain. - [~] Implement real rate-distortion selection and make quality and effort change work, size, and output quality. The complete luma and joint chroma candidate sets, including chroma-from-luma and filter-intra, now perform live rate-distortion selection; quality mapping, effort-dependent pruning, and the remaining searches are not implemented. @@ -858,6 +858,7 @@ Encoder verification contract: - [~] Live luma palette selection now follows current libaom's dominant-color and one-dimensional K-means candidate families, cache-bias threshold, sorted duplicate removal, active-edge map extension, and strict winner tie order. It improves on speed-configured libaom by evaluating both candidate families at every legal 2-through-8 size without early pruning, then exhaustively evaluates every legal transform using the existing SIMD prediction, residual, transform, quantization, and reconstruction operators. Candidate storage remains bounded stack memory; the reusable 32 KiB map owner is allocated only when an eligible block enters palette search. A production tile test proves that full 8x8 and clipped 5x3 blocks at 8 and 12 bits select exact colors and indices, extend the visible edges through coded padding, reconstruct every sample without coefficients, and emit a nonempty tile. The complete 57-case intra-superblock set, 8,931-case AVIF set, and 230-case HEIF set pass direct foreground net11 Release VSTest. The exact Release test-project build reports 1,992 baseline warnings and zero errors; Roslynk reports zero compiler errors and no touched-file analyzer warnings. Production frame activation remains gated until chroma palette mode and its rate accounting are complete. - [~] Paired chroma palette clustering now preserves current libaom's squared two-component distance, first-centroid tie order, independently rounded U/V means, paired deterministic empty-cluster replacement, preceding-state retention on increased distortion, and 50-iteration limit. Keeping the source planes separate avoids interleave/deinterleave copies and improves on libaom's AVX2 ceiling with Vector512, Vector256, Vector128, then scalar dispatch through ImageSharp's shared vector-count helpers. Three independent tests cover exact paired convergence, midpoint initialization, 12-bit distance and index parity, untouched destination bounds, and every intrinsic tier. The exact Release test-project build reports 1,992 baseline warnings and zero errors; the focused three-case set, complete 8,934-case AVIF set, and complete 230-case HEIF set pass direct foreground net11 Release VSTest. Roslynk reports zero compiler errors and no touched-file analyzer warnings. Candidate integration and production activation remain in the open chroma-palette checkpoint. - [~] Live paired chroma palette selection now follows current libaom's complete 2-through-8 color-size search, U-plane neighbor-cache snapping, stable U-ordered color pairs, shared U/V index map, implicit DCT-DCT transform, and strict rate-distortion winner replacement. It improves on speed-configured libaom by applying no early header-cost pruning, keeps planar U/V source data separate, and reuses the SIMD-first prediction, residual, transform, quantization, and reconstruction operators without allocator-backed candidate storage. The production tile regression proves both palette-mode probability branches, exact paired colors and indices, coefficient-free reconstruction, and nonempty syntax. The complete 58-case intra-superblock set, 8,935-case AVIF set, and 230-case HEIF set pass direct foreground net11 Release VSTest. The exact Release test-project build reports 1,992 baseline warnings and zero errors; Roslynk reports zero compiler errors and no touched-file analyzer warnings. Production frame activation remains the next checkpoint. +- [~] Production palette activation now matches current libaom's default good-quality screen detector: it scans only complete 16x16 luma blocks, normalizes high-bit-depth samples to eight bits, admits 2-through-4-color blocks, and uses the reference's strict greater-than-ten-percent frame-area threshold. A 256-bit stack bitset and a fifth-color early exit replace libaom's larger per-block histogram without changing the decision, allocation, or source precision. The adaptive sequence flag remains enabled, the frame flag is set before picture-state allocation, and intra-block copy remains disabled. Focused regressions prove strict-threshold equality, high-bit-depth normalization, five-color rejection, emitted frame-header activation, production decode, and generated payload retention. The exact Release test-project build reports 1,992 baseline warnings and zero errors; all 8,935 AVIF cases and all 230 HEIF cases pass direct foreground net11 Release VSTest. Current-main `aomdec` at `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727` accepts all 30 regenerated production payloads, including the 54-byte palette case. Roslynk reports zero compiler errors and no touched-file analyzer warnings. - [x] The expanded checkpoint exposed a pre-existing transform-block test that asserted uninitialized pooled padding was zero. The test now initializes the complete physical luma plane with a sentinel and proves the block operation leaves both adjacent padding samples unchanged. The exact net11 Release rebuild remains at 1,005 baseline warnings and zero errors, the focused allocator-order set passes 30 of 30 cases, and the complete HEIF/AV1 namespace passes 8,859 of 8,859 direct VSTest cases with zero failures or skips. - [x] Combined-frame OBU output now counts the byte-aligned frame and tile-group headers, non-final tile-size fields, and owned tile payloads before emitting the OBU size. It retains only the small allocator-owned header scratch and writes each entropy-coded tile span directly from its detached owner, removing the second file-sized allocator rent and complete-payload copy. A 64 KiB regression proves exactly one sub-payload-sized byte rent with a balanced return and verifies the exact streamed tile tail; the existing two-tile round trip proves size-prefix and ordering parity. The focused writer and production-frame set passes 32 of 32 direct net11 VSTest cases, current-main `aomdec` accepts all 29 generated native-format payloads, and the complete HEIF/AV1 namespace passes 8,860 of 8,860 cases with zero failures or skips. - [x] Finalized fixed-block decisions now set the block-level transform-skip flag only when every retained luma and coded chroma transform has zero EOB, matching current libaom's conjunction of per-plane skip state. The previous always-false flag produced legal but redundant non-skip and zero-coefficient syntax. Monochrome and 4:2:0 regressions prove both branches from actual coefficient state; the focused decision and production-frame set passes 32 of 32 direct net11 VSTest cases. Current-main `aomdec` accepts all 29 regenerated payloads, the recorded decoded-frame MD5s are unchanged, and affected 16x16 constant 8-bit and 10-bit payloads are one byte smaller. The complete HEIF/AV1 namespace passes 8,862 of 8,862 cases with zero failures or skips. diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1FrameEncoder.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1FrameEncoder.cs index c803b4526c..6e3630c52a 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1FrameEncoder.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1FrameEncoder.cs @@ -235,6 +235,8 @@ internal static class Av1FrameEncoder where TPixel : unmanaged, IPixel { PrepareSource(configuration, image, source.Frame, sequenceHeader.ColorConfig); + frameHeader.AllowScreenContentTools = Av1ScreenContentDetector.IsPaletteLikely(source.Frame); + using Av1EncoderPictureBuffer picture = new( configuration, sequenceHeader, @@ -276,6 +278,8 @@ internal static class Av1FrameEncoder where TPixel : unmanaged, IPixel { PrepareSource(configuration, image, source.Frame, sequenceHeader.ColorConfig); + frameHeader.AllowScreenContentTools = Av1ScreenContentDetector.IsPaletteLikely(source.Frame); + using Av1EncoderPictureBuffer picture = new( configuration, sequenceHeader, diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1ScreenContentDetector.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1ScreenContentDetector.cs new file mode 100644 index 0000000000..919f5557a7 --- /dev/null +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1ScreenContentDetector.cs @@ -0,0 +1,122 @@ +// Copyright (c) Six Labors. +// Licensed under the Six Labors Split License. + +namespace SixLabors.ImageSharp.Formats.Heif.Av1.Pipeline; + +/// +/// Detects frames that can benefit from AV1 screen-content coding tools. +/// +internal static class Av1ScreenContentDetector +{ + private const int DetectionBlockLength = 16; + private const int DetectionBlockArea = DetectionBlockLength * DetectionBlockLength; + private const int MaximumPaletteColorCount = 4; + + /// + /// Converts native source samples to the eight-bit domain used by screen-content detection. + /// + /// The native sample storage type. + private interface ISampleOperator + where TSample : unmanaged + { + /// + /// Converts one native sample to eight-bit precision. + /// + /// The source sample. + /// The number of low bits removed from high-bit-depth samples. + /// The normalized sample. + public static abstract int ToEightBit(TSample value, int bitDepthShift); + } + + /// + /// Detects palette-friendly content in an eight-bit source frame. + /// + /// The converted source frame. + /// when palette tools should be enabled; otherwise, . + public static bool IsPaletteLikely(Av1EncoderFrame source) + => IsPaletteLikely(source); + + /// + /// Detects palette-friendly content in a high-bit-depth source frame. + /// + /// The converted source frame. + /// when palette tools should be enabled; otherwise, . + public static bool IsPaletteLikely(Av1EncoderFrame source) + => IsPaletteLikely(source); + + private static bool IsPaletteLikely(Av1EncoderFrame source) + where TSample : unmanaged + where TOperator : struct, ISampleOperator + { + Av1EncoderFrame.PlanarView view = source.View; + int width = source.Width; + int height = source.Height; + long frameArea = (long)width * height; + int bitDepthShift = source.LumaBitDepth - 8; + int paletteBlockCount = 0; + Span seenColors = stackalloc ulong[4]; + + // Complete 16x16 blocks and the strict frame-area threshold preserve the reference detector's decision. + for (int blockRow = 0; blockRow + DetectionBlockLength <= height; blockRow += DetectionBlockLength) + { + for (int blockColumn = 0; blockColumn + DetectionBlockLength <= width; blockColumn += DetectionBlockLength) + { + seenColors.Clear(); + int colorCount = 0; + for (int row = 0; row < DetectionBlockLength && colorCount <= MaximumPaletteColorCount; row++) + { + ReadOnlySpan samples = view + .GetLumaRowSpan(blockRow + row) + .Slice(blockColumn, DetectionBlockLength); + + // Histogram updates depend on each sample value, so a compact scalar bitset avoids gather/scatter overhead. + for (int column = 0; column < samples.Length; column++) + { + int value = TOperator.ToEightBit(samples[column], bitDepthShift); + int wordIndex = value >> 6; + ulong mask = 1UL << (value & 63); + ref ulong word = ref seenColors[wordIndex]; + if ((word & mask) == 0) + { + word |= mask; + colorCount++; + if (colorCount > MaximumPaletteColorCount) + { + break; + } + } + } + } + + if (colorCount > 1 && colorCount <= MaximumPaletteColorCount) + { + paletteBlockCount++; + if ((long)paletteBlockCount * DetectionBlockArea * 10 > frameArea) + { + return true; + } + } + } + } + + return false; + } + + /// + /// Preserves native eight-bit samples. + /// + private readonly struct ByteSampleOperator : ISampleOperator + { + /// + public static int ToEightBit(byte value, int bitDepthShift) => value; + } + + /// + /// Normalizes high-bit-depth samples to eight-bit precision. + /// + private readonly struct UShortSampleOperator : ISampleOperator + { + /// + public static int ToEightBit(ushort value, int bitDepthShift) => value >> bitDepthShift; + } +} diff --git a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs index 309a32664d..5757dba8f5 100644 --- a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs +++ b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs @@ -146,6 +146,128 @@ public class Av1EncoderFrameTests } } + [Fact] + public void ScreenContentDetectorMatchesLibaomPaletteThreshold() + { + const int width = 160; + const int height = 16; + using Av1EncoderFrameBuffer byteFrame = new( + Configuration.Default, + width, + height, + 8, + Av1ColorFormat.Yuv400, + 0, + 0); + + for (int row = 0; row < height; row++) + { + Span samples = byteFrame.Frame.View.GetLumaRowSpan(row)[..width]; + samples.Fill(96); + for (int column = 0; column < 16; column++) + { + samples[column] = column < 8 ? (byte)32 : (byte)224; + } + } + + // One qualifying block is exactly ten percent of this frame, and the reference threshold is strict. + Assert.False(Av1ScreenContentDetector.IsPaletteLikely(byteFrame.Frame)); + for (int row = 0; row < height; row++) + { + Span samples = byteFrame.Frame.View.GetLumaRowSpan(row); + for (int column = 16; column < 32; column++) + { + samples[column] = column < 24 ? (byte)48 : (byte)208; + } + } + + Assert.True(Av1ScreenContentDetector.IsPaletteLikely(byteFrame.Frame)); + using Av1EncoderFrameBuffer highBitDepthFrame = new( + Configuration.Default, + 16, + 16, + 10, + Av1ColorFormat.Yuv400, + 0, + 0); + + for (int row = 0; row < 16; row++) + { + Span samples = highBitDepthFrame.Frame.View.GetLumaRowSpan(row); + for (int column = 0; column < 16; column++) + { + samples[column] = column < 8 ? (ushort)128 : (ushort)131; + } + } + + Assert.False(Av1ScreenContentDetector.IsPaletteLikely(highBitDepthFrame.Frame)); + for (int row = 0; row < 16; row++) + { + Span samples = highBitDepthFrame.Frame.View.GetLumaRowSpan(row); + samples[8..16].Fill(640); + } + + Assert.True(Av1ScreenContentDetector.IsPaletteLikely(highBitDepthFrame.Frame)); + for (int row = 0; row < 16; row++) + { + Span samples = highBitDepthFrame.Frame.View.GetLumaRowSpan(row); + for (int column = 0; column < 16; column++) + { + samples[column] = (ushort)((column % 5) * 200); + } + } + + Assert.False(Av1ScreenContentDetector.IsPaletteLikely(highBitDepthFrame.Frame)); + } + + [Fact] + public void EncodeActivatesPaletteToolsForScreenContent() + { + const int width = 16; + const int height = 16; + using Image source = new(width, height); + for (int row = 0; row < height; row++) + { + Span pixels = source.Frames.RootFrame.PixelBuffer.DangerousGetRowSpan(row); + for (int column = 0; column < width; column++) + { + pixels[column] = (((column >> 2) + (row >> 2)) & 1) == 0 + ? new Rgba32(224, 32, 32) + : new Rgba32(32, 32, 224); + } + } + + ObuColorConfig colorConfig = CreateColorConfig(Av1BitDepth.EightBit, Av1ColorFormat.Yuv444); + using MemoryStream stream = new(); + _ = Av1FrameEncoder.Encode( + Configuration.Default, + source.Frames.RootFrame, + stream, + colorConfig, + qIndex: 37); + + byte[] payload = stream.ToArray(); + Av1BitStreamReader reader = new(payload); + Av1TileDecoderStub tileReader = new(); + ObuReader obuReader = new(); + obuReader.ReadAll(ref reader, payload.Length, () => tileReader); + ObuFrameHeader frameHeader = Assert.IsType(obuReader.FrameHeader); + Assert.True(frameHeader.AllowScreenContentTools); + Assert.False(frameHeader.AllowIntraBlockCopy); + using Av1Decoder decoder = new(Configuration.Default); + using Image decoded = decoder.Decode(payload); + Assert.Equal(new Size(width, height), decoded.Size); + + string outputDirectory = Path.Combine( + TestEnvironment.ActualOutputDirectoryFullPath, + "Formats", + "Heif", + "Av1"); + + Directory.CreateDirectory(outputDirectory); + File.WriteAllBytes(Path.Combine(outputDirectory, "encoder-frame-16x16-8b-444-palette.obu"), payload); + } + [Fact] public void PrepareSourceConvertsRgba32DirectlyIntoBorderedEightBitPlane() {