diff --git a/HEIF_IMPLEMENTATION_PLAN.md b/HEIF_IMPLEMENTATION_PLAN.md index 2886165b5e..3aae2daf4a 100644 --- a/HEIF_IMPLEMENTATION_PLAN.md +++ b/HEIF_IMPLEMENTATION_PLAN.md @@ -825,11 +825,11 @@ Encoder verification contract: - [~] A non-owning encoder-frame view now separates visible conversion regions from coded regions and performs complete left, top, right, bottom, and corner extension across each bordered plane. Current libaom uses 8-sample-aligned coded dimensions, a 32-sample-aligned luma stride with chroma stride derived from it, and a 64-pixel luma border for non-resized all-intra encoding. One operation-ready frame owner now rents the aligned Y, U, and V storage contiguously, exposes non-owning `Buffer2D` plane views, and returns the rent exactly once. A 4K 4:2:0 frame occupies about 13.0 MiB at 8-bit or 26.0 MiB at 10/12-bit; source and reconstruction therefore remain distinct frame owners rather than adding a full-frame copy. The corrected tests use this real ownership path and verify the exact 54 KiB 64x64 4:2:0 rent. The frame-encoder operation now instantiates matching source and reconstruction owners with ordinary `using` lifetimes and converts packed pixels directly into the source owner before extension. - [~] Temporal delimiter, sequence header, frame header, combined-frame tile-group writing, and an internal reduced-still-picture frame operation now exist locally. The remaining required metadata, padding, multi-tile, option, and public encoder paths are not complete. - [~] Implement superblock and partition analysis for every permitted block size and partition. The current baseline deliberately splits every in-frame node to 8x8 blocks and records decisions in current-libaom writer preorder; block-size selection and non-split partition analysis remain. -- [~] Implement intra mode search, palette, filter intra, chroma-from-luma, and intra-block copy decisions. Live luma search now covers all 13 zero-angle base modes and all six nonzero adjustments for each of the eight directional modes. Joint spatial chroma search covers the same 61 candidates, combines both chroma planes in one rate-distortion decision, and preserves the winning shared angle adjustment. Chroma-from-luma now searches the complete signed alpha alphabet from reconstructed luma and retains its joint U/V syntax. Filter-intra now searches all five predictors after ordinary luma modes. Palette entropy, retained state, production syntax, and exhaustive luma palette selection are complete; chroma palette and intra-block-copy mode decisions remain. +- [~] Implement intra mode search, palette, filter intra, chroma-from-luma, and intra-block copy decisions. Live luma search now covers all 13 zero-angle base modes and all six nonzero adjustments for each of the eight directional modes. Joint spatial chroma search covers the same 61 candidates, combines both chroma planes in one rate-distortion decision, and preserves the winning shared angle adjustment. Chroma-from-luma now searches the complete signed alpha alphabet from reconstructed luma and retains its joint U/V syntax. Filter-intra now searches all five predictors after ordinary luma modes. Palette entropy, retained state, production syntax, and exhaustive luma and paired chroma palette selection are complete; intra-block-copy mode decisions remain. - [ ] Implement inter mode search for bounded sequences, including reference selection and the decoder-supported inter tools. - [~] Current-libaom `av1_quantize_fp_no_qmatrix` arithmetic is implemented as a closed generic forward-quantizer family with Vector512, Vector256, Vector128, and scalar paths, raster-order output, coded 64-point coefficient limits, and scan-order EOB selection. Transform search, coefficient optimization, and lossless behavior remain. - [~] Implement real rate-distortion selection and make quality and effort change work, size, and output quality. The complete luma and joint chroma candidate sets, including chroma-from-luma and filter-intra, now perform live rate-distortion selection; quality mapping, effort-dependent pruning, and the remaining searches are not implemented. -- [~] Encoder rate accounting converts the entropy writer's live inverse cumulative distributions into current-libaom fixed-point symbol costs without allocating or duplicating probability state. Read-only luma-mode, directional-delta, filter-intra, chroma-mode, block-skip, transform-size, transform-block-skip, and complete transform-coefficient queries share the exact distributions mutated by the subsequent entropy write. Complete coefficient costing follows current libaom's optimized shape: it returns immediately for an empty transform, uses the EOB-specific base-range context, fuses magnitude, sign, base-range, and Golomb accounting into one reverse traversal, and combines repeated full base-range chunks instead of replaying each emitted symbol. Tile-lifetime level and context scratch is reused, the one-coefficient path neither clears nor initializes the forward-neighbor level map, and steady-state queries allocate nothing. Transform-size writing and costing share one subdivision-depth calculation, while shared closed symbol operations keep the writer and cost mappings for transform skip, transform type, and EOB syntax identical without forcing the estimator through the writer's slower two-pass coefficient traversal. The current-libaom fixed-point RD combiner preserves 64-bit distortion and rounds the weighted 1/512-bit rate at the required boundary. Its key-frame multiplier follows libaom's squared DC-quantizer formula and exact 10/12-bit normalization. Live final-block selection evaluates all 61 legal 8x8 luma candidates: the 13 zero-angle base modes in current-libaom order, followed by six nonzero adjustments for each directional mode. Joint chroma selection evaluates the equivalent 61 spatial candidates, combines U and V distortion plus coefficient rate, and charges one live chroma-mode and shared-angle symbol over the actual subsampled 4x4, 4x8, or 8x8 geometry. Chroma-from-luma subsamples the reconstructed luma block once into fixed-stride Q3 stack scratch, subtracts the rounded mean, evaluates all 33 signed alpha values independently for each plane with complete transform RD, and combines the cached plane results across all 1,088 valid joint pairs with one live sign cost and the conditional U/V magnitude costs. This is the allocation-free equivalent of current libaom's exhaustive 33-value path: it requires 66 evaluation transforms rather than transforming every joint pair, preserves DC-before-CfL-before-spatial tie order, and fixes the implicit chroma transform to DCT-DCT. Filter-intra follows ordinary luma candidates, searches all five predictors in syntax order, and evaluates every legal transform while reusing one prepared prediction and source residual per filter mode. Every candidate includes its live mode, angle, filter mode, alpha, and coefficient rate plus normalized pixel-domain distortion. Each prepared reference edge retains the common-corner prefix and twice the transform dimension required by directional prediction. A shared encoder/decoder availability calculation selects reconstructed top-right and bottom-left extensions according to tile, frame, superblock, and block reconstruction order; unavailable extensions repeat the nearest coded endpoint. Missing top or left edges retain current libaom's perpendicular-sample and bit-depth-midpoint rules. Directional prediction applies the AV1 three-degree adjustment step and reuses transform workspace for zone-three transposition before the transform overwrites it, keeping candidate evaluation allocation-free. The winning luma and chroma signed adjustments are retained in the packed final-block state consumed by the tile writer. The tile writer invokes these stack-only selectors after mapping current neighbors and immediately before writing each block, so later decisions see reconstructed samples, coefficient contexts, and CDF updates from every preceding block. Block skip is read only after the callback has combined every coded plane. Luma candidate scratch remains one 8x8 reconstruction and one 8x8 coefficient span on the stack; chroma uses one transform-sized reconstruction and coefficient span for each of U and V. Only a newly winning candidate is copied into retained frame storage. Production fixtures force every luma base predictor, both extreme adjustments in all three directional zones, available top-right and bottom-left extensions, high-bit-depth adjustment propagation, exact signed luma and chroma angle-rate terms, joint U/V decisions, packed chroma state, and 4:2:0, 4:2:2, and 4:4:4 transform geometry. The CfL fixtures derive target chroma from a pilot production encode's actual reconstructed luma through an independent scalar Q3 oracle and prove exact positive/negative alpha syntax plus zero-residual DCT-DCT reconstruction for all three subsampling geometries at 8, 10, and 12 bits. The stable fixed-DC traversal comparison uses neutral samples for which both the baseline and live search are contractually DC and skipped, instead of relying on textured content to happen to select the baseline mode. Luma palette selection now evaluates dominant-color and one-dimensional K-means candidates for every legal size, snaps near-cache colors with the reference threshold and tie order, removes duplicate snapped colors, extends boundary maps from active samples, and performs complete transform rate-distortion search. Ordinary DC and filter-intra candidates pay the palette-disabled symbol whenever screen-content syntax is enabled. The exact net11 Release rebuild reports 1,992 test-project warnings and zero errors, all 57 intra-superblock cases pass, all 8,931 AVIF cases pass, and all 230 HEIF cases pass. Remaining mode decision work includes transform-size search, broader joint mode/transform refinement, chroma palette search, partition search, and effort-dependent pruning. Non-empty intra blocks deliberately remain non-skipped, matching current libaom; later inter mode selection owns its distinct skip-transform RD decision. +- [~] Encoder rate accounting converts the entropy writer's live inverse cumulative distributions into current-libaom fixed-point symbol costs without allocating or duplicating probability state. Read-only luma-mode, directional-delta, filter-intra, chroma-mode, block-skip, transform-size, transform-block-skip, and complete transform-coefficient queries share the exact distributions mutated by the subsequent entropy write. Complete coefficient costing follows current libaom's optimized shape: it returns immediately for an empty transform, uses the EOB-specific base-range context, fuses magnitude, sign, base-range, and Golomb accounting into one reverse traversal, and combines repeated full base-range chunks instead of replaying each emitted symbol. Tile-lifetime level and context scratch is reused, the one-coefficient path neither clears nor initializes the forward-neighbor level map, and steady-state queries allocate nothing. Transform-size writing and costing share one subdivision-depth calculation, while shared closed symbol operations keep the writer and cost mappings for transform skip, transform type, and EOB syntax identical without forcing the estimator through the writer's slower two-pass coefficient traversal. The current-libaom fixed-point RD combiner preserves 64-bit distortion and rounds the weighted 1/512-bit rate at the required boundary. Its key-frame multiplier follows libaom's squared DC-quantizer formula and exact 10/12-bit normalization. Live final-block selection evaluates all 61 legal 8x8 luma candidates: the 13 zero-angle base modes in current-libaom order, followed by six nonzero adjustments for each directional mode. Joint chroma selection evaluates the equivalent 61 spatial candidates, combines U and V distortion plus coefficient rate, and charges one live chroma-mode and shared-angle symbol over the actual subsampled 4x4, 4x8, or 8x8 geometry. Chroma-from-luma subsamples the reconstructed luma block once into fixed-stride Q3 stack scratch, subtracts the rounded mean, evaluates all 33 signed alpha values independently for each plane with complete transform RD, and combines the cached plane results across all 1,088 valid joint pairs with one live sign cost and the conditional U/V magnitude costs. This is the allocation-free equivalent of current libaom's exhaustive 33-value path: it requires 66 evaluation transforms rather than transforming every joint pair, preserves DC-before-CfL-before-spatial tie order, and fixes the implicit chroma transform to DCT-DCT. Filter-intra follows ordinary luma candidates, searches all five predictors in syntax order, and evaluates every legal transform while reusing one prepared prediction and source residual per filter mode. Every candidate includes its live mode, angle, filter mode, alpha, and coefficient rate plus normalized pixel-domain distortion. Each prepared reference edge retains the common-corner prefix and twice the transform dimension required by directional prediction. A shared encoder/decoder availability calculation selects reconstructed top-right and bottom-left extensions according to tile, frame, superblock, and block reconstruction order; unavailable extensions repeat the nearest coded endpoint. Missing top or left edges retain current libaom's perpendicular-sample and bit-depth-midpoint rules. Directional prediction applies the AV1 three-degree adjustment step and reuses transform workspace for zone-three transposition before the transform overwrites it, keeping candidate evaluation allocation-free. The winning luma and chroma signed adjustments are retained in the packed final-block state consumed by the tile writer. The tile writer invokes these stack-only selectors after mapping current neighbors and immediately before writing each block, so later decisions see reconstructed samples, coefficient contexts, and CDF updates from every preceding block. Block skip is read only after the callback has combined every coded plane. Luma candidate scratch remains one 8x8 reconstruction and one 8x8 coefficient span on the stack; chroma uses one transform-sized reconstruction and coefficient span for each of U and V. Only a newly winning candidate is copied into retained frame storage. Production fixtures force every luma base predictor, both extreme adjustments in all three directional zones, available top-right and bottom-left extensions, high-bit-depth adjustment propagation, exact signed luma and chroma angle-rate terms, joint U/V decisions, packed chroma state, and 4:2:0, 4:2:2, and 4:4:4 transform geometry. The CfL fixtures derive target chroma from a pilot production encode's actual reconstructed luma through an independent scalar Q3 oracle and prove exact positive/negative alpha syntax plus zero-residual DCT-DCT reconstruction for all three subsampling geometries at 8, 10, and 12 bits. The stable fixed-DC traversal comparison uses neutral samples for which both the baseline and live search are contractually DC and skipped, instead of relying on textured content to happen to select the baseline mode. Luma palette selection now evaluates dominant-color and one-dimensional K-means candidates for every legal size, snaps near-cache colors with the reference threshold and tie order, removes duplicate snapped colors, extends boundary maps from active samples, and performs complete transform rate-distortion search. Ordinary DC and filter-intra candidates pay the palette-disabled symbol whenever screen-content syntax is enabled. The exact net11 Release rebuild reports 1,992 test-project warnings and zero errors, all 58 intra-superblock cases pass, all 8,935 AVIF cases pass, and all 230 HEIF cases pass. Remaining mode decision work includes transform-size search, broader joint mode/transform refinement, partition search, and effort-dependent pruning. Non-empty intra blocks deliberately remain non-skipped, matching current libaom; later inter mode selection owns its distinct skip-transform RD decision. - [~] The tile writer now publishes one packed coefficient context per covered 4x4 edge unit and derives luma/chroma skip plus DC-sign contexts from the complete transform edges using current-libaom units. Partition, transform, and coefficient neighbor state retains only the above and left context regions used by current libaom; the unused third top-left region, its granularity state, and its unused sentinel are removed. One picture owner now packs segmentation plus every tile's partition, luma, chroma, and transform edges into one clean byte allocation with typed non-owning views; together with the separately typed packed mode-information owner, the complete picture state uses two allocator rents rather than seven. Exact aligned lengths, clean initialization, and balanced exactly-once returns are covered in Release. Multi-tile payload ownership and verified CDF update behavior remain. - [~] Encoder mode information now uses a frame-owned integer alias grid over a packed 8-byte value allocation, matching current libaom's `mi_grid_base` and `mi_alloc` relationship without a managed object or reference per 4x4 entry. The visible dimensions are aligned to eight luma samples, the grid stride and allocated row count are aligned to 32 mode-information units, and optional 8x8 allocation granularity reduces the value store in both dimensions exactly as current libaom does. One clean ImageSharp byte owner contains both independently typed regions, reducing libaom's two allocation lifetimes to one without a copy. At 4K, the 4x4 layout occupies about 6.0 MiB in total; the 8x8 layout occupies about 3.0 MiB. Exact geometry, clean allocation, typed lengths, aligned mapping, untouched row padding, and exactly-once return pass 4 of 4 direct net11 VSTest cases in Release. Every coded 4x4 cell covered by square, rectangular, or clipped edge blocks maps to its owning allocation entry before context-dependent symbols are written. Packed syntax, relative neighbor lookup, full block mapping, writer traversal, entropy, and OBU coverage pass 1,947 of 1,947 direct net11 VSTest cases in Release; complete mode decision still remains. - [~] The final-block decision workspace uses one reusable 8.3 KiB ImageSharp allocator owner. It contains 1,024 explicitly packed 8-byte final-block entries and the 341 preorder partition bytes required by a complete 128x128-through-8x8 quadtree, replacing separate managed arrays. Palette colors now have their own current-block value and are copied only to the picture edges that later blocks can reference, so enabling palette mode does not add 50 bytes to every final-block entry. Construction and the explicit per-superblock reset initialize every syntax field, including the nonzero sentinel that disables filter-intra prediction; pooled quantizer, prediction, partition, and current-palette bytes cannot leak into the next decision pass. Complete mode decision still remains. @@ -857,6 +857,7 @@ Encoder verification contract: - [~] Luma palette clustering now follows current libaom's one-dimensional search primitive exactly: equal-interval midpoint initialization, first-color tie order, rounded centroid means, deterministic empty-cluster replacement, the 50-iteration limit, and retention of the preceding state when distortion increases. Nearest-color assignment improves on libaom's AVX2 implementation by dispatching Vector512, Vector256, Vector128, then scalar through ImageSharp's shared vector-count helpers. The primitive uses only bounded stack scratch and introduces no allocator rent, managed array, or per-row copy. Three independent tests cover exact centroid convergence, initialization order, 12-bit nearest-color distortion, destination bounds, and every hardware-intrinsic tier. The complete AVIF set passes 8,930 of 8,930 cases and the HEIF set passes 230 of 230 cases through direct foreground net11 Release VSTest. The exact net11 Release rebuild reports 1,050 solution warnings and zero errors, and Roslynk reports zero compiler errors. Candidate enumeration, palette-cache snapping, transform RD selection, and production activation remain in the open luma-palette checkpoint. - [~] Live luma palette selection now follows current libaom's dominant-color and one-dimensional K-means candidate families, cache-bias threshold, sorted duplicate removal, active-edge map extension, and strict winner tie order. It improves on speed-configured libaom by evaluating both candidate families at every legal 2-through-8 size without early pruning, then exhaustively evaluates every legal transform using the existing SIMD prediction, residual, transform, quantization, and reconstruction operators. Candidate storage remains bounded stack memory; the reusable 32 KiB map owner is allocated only when an eligible block enters palette search. A production tile test proves that full 8x8 and clipped 5x3 blocks at 8 and 12 bits select exact colors and indices, extend the visible edges through coded padding, reconstruct every sample without coefficients, and emit a nonempty tile. The complete 57-case intra-superblock set, 8,931-case AVIF set, and 230-case HEIF set pass direct foreground net11 Release VSTest. The exact Release test-project build reports 1,992 baseline warnings and zero errors; Roslynk reports zero compiler errors and no touched-file analyzer warnings. Production frame activation remains gated until chroma palette mode and its rate accounting are complete. - [~] Paired chroma palette clustering now preserves current libaom's squared two-component distance, first-centroid tie order, independently rounded U/V means, paired deterministic empty-cluster replacement, preceding-state retention on increased distortion, and 50-iteration limit. Keeping the source planes separate avoids interleave/deinterleave copies and improves on libaom's AVX2 ceiling with Vector512, Vector256, Vector128, then scalar dispatch through ImageSharp's shared vector-count helpers. Three independent tests cover exact paired convergence, midpoint initialization, 12-bit distance and index parity, untouched destination bounds, and every intrinsic tier. The exact Release test-project build reports 1,992 baseline warnings and zero errors; the focused three-case set, complete 8,934-case AVIF set, and complete 230-case HEIF set pass direct foreground net11 Release VSTest. Roslynk reports zero compiler errors and no touched-file analyzer warnings. Candidate integration and production activation remain in the open chroma-palette checkpoint. +- [~] Live paired chroma palette selection now follows current libaom's complete 2-through-8 color-size search, U-plane neighbor-cache snapping, stable U-ordered color pairs, shared U/V index map, implicit DCT-DCT transform, and strict rate-distortion winner replacement. It improves on speed-configured libaom by applying no early header-cost pruning, keeps planar U/V source data separate, and reuses the SIMD-first prediction, residual, transform, quantization, and reconstruction operators without allocator-backed candidate storage. The production tile regression proves both palette-mode probability branches, exact paired colors and indices, coefficient-free reconstruction, and nonempty syntax. The complete 58-case intra-superblock set, 8,935-case AVIF set, and 230-case HEIF set pass direct foreground net11 Release VSTest. The exact Release test-project build reports 1,992 baseline warnings and zero errors; Roslynk reports zero compiler errors and no touched-file analyzer warnings. Production frame activation remains the next checkpoint. - [x] The expanded checkpoint exposed a pre-existing transform-block test that asserted uninitialized pooled padding was zero. The test now initializes the complete physical luma plane with a sentinel and proves the block operation leaves both adjacent padding samples unchanged. The exact net11 Release rebuild remains at 1,005 baseline warnings and zero errors, the focused allocator-order set passes 30 of 30 cases, and the complete HEIF/AV1 namespace passes 8,859 of 8,859 direct VSTest cases with zero failures or skips. - [x] Combined-frame OBU output now counts the byte-aligned frame and tile-group headers, non-final tile-size fields, and owned tile payloads before emitting the OBU size. It retains only the small allocator-owned header scratch and writes each entropy-coded tile span directly from its detached owner, removing the second file-sized allocator rent and complete-payload copy. A 64 KiB regression proves exactly one sub-payload-sized byte rent with a balanced return and verifies the exact streamed tile tail; the existing two-tile round trip proves size-prefix and ordering parity. The focused writer and production-frame set passes 32 of 32 direct net11 VSTest cases, current-main `aomdec` accepts all 29 generated native-format payloads, and the complete HEIF/AV1 namespace passes 8,860 of 8,860 cases with zero failures or skips. - [x] Finalized fixed-block decisions now set the block-level transform-skip flag only when every retained luma and coded chroma transform has zero EOB, matching current libaom's conjunction of per-plane skip state. The previous always-false flag produced legal but redundant non-skip and zero-coefficient syntax. Monochrome and 4:2:0 regressions prove both branches from actual coefficient state; the focused decision and production-frame set passes 32 of 32 direct net11 VSTest cases. Current-main `aomdec` accepts all 29 regenerated payloads, the recorded decoded-frame MD5s are unchanged, and affected 16x16 constant 8-bit and 10-bit payloads are one byte smaller. The complete HEIF/AV1 namespace passes 8,862 of 8,862 cases with zero failures or skips. diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs index 0150a5089e..b7991d7540 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs @@ -53,6 +53,7 @@ internal static partial class Av1IntraSuperblockEncoder Span retainedRedCoefficients, ref Av1EncoderTransformBlockState retainedBlueState, ref Av1EncoderTransformBlockState retainedRedState, + ref Av1EncoderPaletteInfo paletteInfo, out int selectedAngleDelta, out byte selectedChromaFromLumaIndex, out sbyte selectedChromaFromLumaSigns) @@ -163,6 +164,10 @@ internal static partial class Av1IntraSuperblockEncoder int deltaCount = AngleDeltaSearchOrder.Length; int directionalModeCount = (int)Av1ChromaPredictionMode.Directional67Degrees - (int)Av1ChromaPredictionMode.Vertical + 1; int candidateCount = baseModeCount + (directionalModeCount * deltaCount); + bool hasLumaPalette = paletteInfo.PaletteSizes[0] != 0; + int paletteDisabledCost = this.picture.Parent.FrameHeader.AllowScreenContentTools + ? writer.GetPaletteUvModeCost(false, hasLumaPalette) + : 0; // Spatial base modes precede the six nonzero adjustments for each directional mode. // Chroma-from-luma remains a separate search because it consumes reconstructed luma AC state. @@ -202,6 +207,7 @@ internal static partial class Av1IntraSuperblockEncoder hasAbove, blueContext, redContext, + paletteDisabledCost, candidateBlueReconstruction[..sampleCount], candidateRedReconstruction[..sampleCount], candidateBlueCoefficients[..sampleCount], @@ -434,6 +440,35 @@ internal static partial class Av1IntraSuperblockEncoder } } + if (this.picture.Parent.FrameHeader.AllowScreenContentTools && + this.SelectChromaPalette( + writer, + macroBlock, + modeInfo, + lumaOrigin, + chromaOrigin, + tileIndex, + lumaMode, + transformSize, + blueContext, + redContext, + candidateBlueReconstruction[..sampleCount], + candidateRedReconstruction[..sampleCount], + candidateBlueCoefficients[..sampleCount], + candidateRedCoefficients[..sampleCount], + retainedBlueCoefficients, + retainedRedCoefficients, + ref retainedBlueState, + ref retainedRedState, + ref bestCost, + ref paletteInfo)) + { + bestMode = Av1ChromaPredictionMode.DC; + selectedAngleDelta = 0; + selectedChromaFromLumaIndex = 0; + selectedChromaFromLumaSigns = 0; + } + return bestMode; } @@ -502,6 +537,7 @@ internal static partial class Av1IntraSuperblockEncoder bool hasAbove, Av1TransformBlockContext blueContext, Av1TransformBlockContext redContext, + int paletteDisabledCost, Span candidateBlueReconstruction, Span candidateRedReconstruction, Span candidateBlueCoefficients, @@ -568,6 +604,11 @@ internal static partial class Av1IntraSuperblockEncoder chromaMode, angleDelta); + if (chromaMode == Av1ChromaPredictionMode.DC) + { + rate += paletteDisabledCost; + } + rate += writer.GetCoefficientCost( transformSize, transformType, diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaPaletteModeDecision.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaPaletteModeDecision.cs new file mode 100644 index 0000000000..69f67f2dad --- /dev/null +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaPaletteModeDecision.cs @@ -0,0 +1,373 @@ +// Copyright (c) Six Labors. +// Licensed under the Six Labors Split License. + +using SixLabors.ImageSharp.Formats.Heif.Av1.Entropy; +using SixLabors.ImageSharp.Formats.Heif.Av1.OpenBitstreamUnit; +using SixLabors.ImageSharp.Formats.Heif.Av1.Prediction; +using SixLabors.ImageSharp.Formats.Heif.Av1.Tiling; +using SixLabors.ImageSharp.Formats.Heif.Av1.Transform; +using SixLabors.ImageSharp.Memory; + +namespace SixLabors.ImageSharp.Formats.Heif.Av1.Pipeline; + +/// +/// Provides paired chroma palette mode decisions for intra encoding. +/// +internal static partial class Av1IntraSuperblockEncoder +{ + internal partial struct ModeDecision + where TSample : unmanaged + where TOperator : struct, IBlockEncodingOperator + { + private bool SelectChromaPalette( + Av1SymbolEncoder writer, + Av1MacroBlockD macroBlock, + Av1MacroBlockModeInfo modeInfo, + Point lumaOrigin, + Point chromaOrigin, + ushort tileIndex, + Av1PredictionMode lumaMode, + Av1TransformSize transformSize, + Av1TransformBlockContext blueContext, + Av1TransformBlockContext redContext, + Span candidateBlueReconstruction, + Span candidateRedReconstruction, + Span candidateBlueCoefficients, + Span candidateRedCoefficients, + Span retainedBlueCoefficients, + Span retainedRedCoefficients, + ref Av1EncoderTransformBlockState retainedBlueState, + ref Av1EncoderTransformBlockState retainedRedState, + ref long bestCost, + ref Av1EncoderPaletteInfo paletteInfo) + { + const Av1BlockSize BlockSize = Av1BlockSize.Block8x8; + const int LumaBlockLength = 8; + const int MaximumSampleCount = LumaBlockLength * LumaBlockLength; + ObuColorConfig colorConfig = this.picture.Sequence.SequenceHeader.ColorConfig; + int subsamplingX = colorConfig.SubSamplingX ? 1 : 0; + int subsamplingY = colorConfig.SubSamplingY ? 1 : 0; + int width = transformSize.GetWidth(); + int height = transformSize.GetHeight(); + ObuFrameSize frameSize = this.picture.Parent.FrameHeader.FrameSize; + int rows = Math.Min(LumaBlockLength, frameSize.FrameHeight - lumaOrigin.Y) >> subsamplingY; + int columns = Math.Min(LumaBlockLength, frameSize.FrameWidth - lumaOrigin.X) >> subsamplingX; + int activeSampleCount = rows * columns; + Span blueSamples = stackalloc short[MaximumSampleCount]; + Span redSamples = stackalloc short[MaximumSampleCount]; + blueSamples = blueSamples[..activeSampleCount]; + redSamples = redSamples[..activeSampleCount]; + Buffer2DRegion blueSource = this.source.GetPlane(Av1Plane.U); + Buffer2DRegion redSource = this.source.GetPlane(Av1Plane.V); + TOperator.CopyPaletteSamples(blueSource, chromaOrigin, rows, columns, blueSamples); + TOperator.CopyPaletteSamples(redSource, chromaOrigin, rows, columns, redSamples); + + Span uniqueBlueColors = stackalloc short[MaximumSampleCount]; + Span uniqueRedColors = stackalloc short[MaximumSampleCount]; + int uniqueBlueColorCount = 0; + int uniqueRedColorCount = 0; + short blueMinimum = blueSamples[0]; + short blueMaximum = blueSamples[0]; + short redMinimum = redSamples[0]; + short redMaximum = redSamples[0]; + for (int sampleIndex = 0; sampleIndex < activeSampleCount; sampleIndex++) + { + short blueSample = blueSamples[sampleIndex]; + short redSample = redSamples[sampleIndex]; + if (!uniqueBlueColors[..uniqueBlueColorCount].Contains(blueSample)) + { + uniqueBlueColors[uniqueBlueColorCount++] = blueSample; + } + + if (!uniqueRedColors[..uniqueRedColorCount].Contains(redSample)) + { + uniqueRedColors[uniqueRedColorCount++] = redSample; + } + + blueMinimum = Math.Min(blueMinimum, blueSample); + blueMaximum = Math.Max(blueMaximum, blueSample); + redMinimum = Math.Min(redMinimum, redSample); + redMaximum = Math.Max(redMaximum, redSample); + } + + int maximumColorCount = Math.Max(uniqueBlueColorCount, uniqueRedColorCount); + if (maximumColorCount < 2) + { + return false; + } + + int maximumPaletteSize = Math.Min(maximumColorCount, Av1Constants.PaletteMaxSize); + Av1NeighborArrayUnit paletteContexts = this.picture.PaletteContexts[tileIndex]; + Span colorCache = stackalloc ushort[2 * Av1Constants.PaletteMaxSize]; + int colorCacheSize = Av1TileWriter.GetPaletteCache( + paletteContexts, + macroBlock, + lumaOrigin, + Av1Plane.U, + colorCache); + + colorCache = colorCache[..colorCacheSize]; + int blockSizeContext = Av1TileWriter.GetPaletteBlockSizeContext(BlockSize); + bool hasLumaPalette = paletteInfo.PaletteSizes[0] != 0; + Buffer2DRegion colorIndexMap = this.superblock.Workspace + .GetPaletteMaps() + .GetMap(Av1PlaneType.Uv, width, height); + + Span retainedColorIndexMap = stackalloc byte[MaximumSampleCount]; + Span blueCentroids = stackalloc short[Av1Constants.PaletteMaxSize]; + Span redCentroids = stackalloc short[Av1Constants.PaletteMaxSize]; + Span colorIndices = stackalloc byte[MaximumSampleCount]; + Span bluePrediction = stackalloc TSample[MaximumSampleCount]; + Span redPrediction = stackalloc TSample[MaximumSampleCount]; + Span blueResidual = stackalloc short[MaximumSampleCount]; + Span redResidual = stackalloc short[MaximumSampleCount]; + Span bluePaletteColorStorage = stackalloc ushort[Av1Constants.PaletteMaxSize]; + Span redPaletteColorStorage = stackalloc ushort[Av1Constants.PaletteMaxSize]; + int sampleCount = transformSize.GetSize2d(); + int cacheThreshold = 4 << (this.bitDepth.GetBitCount() - 8); + bool paletteSelected = false; + + // Chroma uses one paired K-means family; exhaustive size search avoids early header-cost pruning. + for (int paletteSize = 2; paletteSize <= maximumPaletteSize; paletteSize++) + { + Span candidateBlueCentroids = blueCentroids[..paletteSize]; + Span candidateRedCentroids = redCentroids[..paletteSize]; + Av1PaletteKMeans2D.InitializeCentroids( + blueMinimum, + blueMaximum, + redMinimum, + redMaximum, + candidateBlueCentroids, + candidateRedCentroids); + + Av1PaletteKMeans2D.Cluster( + blueSamples, + redSamples, + candidateBlueCentroids, + candidateRedCentroids, + colorIndices[..activeSampleCount]); + + for (int colorIndex = 0; colorIndex < paletteSize && !colorCache.IsEmpty; colorIndex++) + { + int minimumDifference = Math.Abs(candidateBlueCentroids[colorIndex] - colorCache[0]); + int nearestCacheIndex = 0; + for (int cacheIndex = 1; cacheIndex < colorCache.Length; cacheIndex++) + { + int difference = Math.Abs(candidateBlueCentroids[colorIndex] - colorCache[cacheIndex]); + if (difference < minimumDifference) + { + minimumDifference = difference; + nearestCacheIndex = cacheIndex; + } + } + + if (minimumDifference <= cacheThreshold) + { + candidateBlueCentroids[colorIndex] = (short)colorCache[nearestCacheIndex]; + } + } + + // U is the coded ordering key, so each swap carries its paired V color with it. + for (int colorIndex = 0; colorIndex < paletteSize - 1; colorIndex++) + { + int minimumIndex = colorIndex; + for (int candidateIndex = colorIndex + 1; candidateIndex < paletteSize; candidateIndex++) + { + if (candidateBlueCentroids[candidateIndex] < candidateBlueCentroids[minimumIndex]) + { + minimumIndex = candidateIndex; + } + } + + if (minimumIndex != colorIndex) + { + (candidateBlueCentroids[colorIndex], candidateBlueCentroids[minimumIndex]) = + (candidateBlueCentroids[minimumIndex], candidateBlueCentroids[colorIndex]); + + (candidateRedCentroids[colorIndex], candidateRedCentroids[minimumIndex]) = + (candidateRedCentroids[minimumIndex], candidateRedCentroids[colorIndex]); + } + } + + Av1PaletteKMeans2D.AssignIndices( + blueSamples, + redSamples, + candidateBlueCentroids, + candidateRedCentroids, + colorIndices); + + for (int row = 0; row < rows; row++) + { + Span mapRow = colorIndexMap.DangerousGetRowSpan(row)[..width]; + colorIndices.Slice(row * columns, columns).CopyTo(mapRow); + mapRow[columns..].Fill(mapRow[columns - 1]); + } + + // The shared U/V map covers the complete declared chroma block even at visible frame edges. + for (int row = rows; row < height; row++) + { + colorIndexMap.DangerousGetRowSpan(rows - 1)[..width] + .CopyTo(colorIndexMap.DangerousGetRowSpan(row)); + } + + Span bluePaletteColors = bluePaletteColorStorage[..paletteSize]; + Span redPaletteColors = redPaletteColorStorage[..paletteSize]; + for (int colorIndex = 0; colorIndex < paletteSize; colorIndex++) + { + bluePaletteColors[colorIndex] = (ushort)candidateBlueCentroids[colorIndex]; + redPaletteColors[colorIndex] = (ushort)candidateRedCentroids[colorIndex]; + } + + TOperator.PreparePalette( + blueSource, + chromaOrigin, + bluePaletteColors, + colorIndexMap, + bluePrediction[..sampleCount], + blueResidual[..sampleCount], + transformSize); + + TOperator.PreparePalette( + redSource, + chromaOrigin, + redPaletteColors, + colorIndexMap, + redPrediction[..sampleCount], + redResidual[..sampleCount], + transformSize); + + Av1EncoderTransformBlockState candidateBlueState = default; + long distortion = TOperator.EncodePredictionCandidate( + this.blockWorkspace, + blueSource, + chromaOrigin, + bluePrediction, + blueResidual, + candidateBlueReconstruction, + candidateBlueCoefficients, + transformSize, + Av1TransformType.DctDct, + Av1Plane.U, + this.quantization.QIndex[0], + this.quantization.DeltaQDc[(int)Av1Plane.U], + this.quantization.DeltaQAc[(int)Av1Plane.U], + this.bitDepth, + ref candidateBlueState); + + Av1EncoderTransformBlockState candidateRedState = default; + distortion += TOperator.EncodePredictionCandidate( + this.blockWorkspace, + redSource, + chromaOrigin, + redPrediction, + redResidual, + candidateRedReconstruction, + candidateRedCoefficients, + transformSize, + Av1TransformType.DctDct, + Av1Plane.V, + this.quantization.QIndex[0], + this.quantization.DeltaQDc[(int)Av1Plane.V], + this.quantization.DeltaQAc[(int)Av1Plane.V], + this.bitDepth, + ref candidateRedState); + + int rate = Av1TileWriter.GetChromaModeCost( + writer, + this.picture.Parent.FrameHeader, + colorConfig, + modeInfo, + BlockSize, + lumaMode, + Av1ChromaPredictionMode.DC, + 0); + + rate += writer.GetPaletteUvModeCost(true, hasLumaPalette); + rate += writer.GetPaletteSizeCost(paletteSize, blockSizeContext, Av1PlaneType.Uv); + rate += Av1SymbolEncoder.GetPaletteUvColorCost( + colorCache, + bluePaletteColors, + redPaletteColors, + this.bitDepth.GetBitCount()); + + rate += writer.GetPaletteColorMapCost( + paletteSize, + Av1PlaneType.Uv, + rows, + columns, + colorIndexMap); + + rate += writer.GetCoefficientCost( + transformSize, + Av1TransformType.DctDct, + lumaMode, + candidateBlueCoefficients, + Av1ComponentType.Chroma, + blueContext, + candidateBlueState.EndOfBlock, + this.picture.Parent.FrameHeader.UseReducedTransformSet, + Av1FilterIntraMode.AllFilterIntraModes); + + rate += writer.GetCoefficientCost( + transformSize, + Av1TransformType.DctDct, + lumaMode, + candidateRedCoefficients, + Av1ComponentType.Chroma, + redContext, + candidateRedState.EndOfBlock, + this.picture.Parent.FrameHeader.UseReducedTransformSet, + Av1FilterIntraMode.AllFilterIntraModes); + + long candidateCost = Av1RateDistortion.GetCost(this.rateMultiplier, rate, distortion); + if (candidateCost < bestCost) + { + Buffer2DRegion blueReconstruction = this.reconstruction.GetPlane(Av1Plane.U); + Buffer2DRegion redReconstruction = this.reconstruction.GetPlane(Av1Plane.V); + CopyCandidate( + candidateBlueReconstruction, + candidateBlueCoefficients, + blueReconstruction, + chromaOrigin, + retainedBlueCoefficients, + transformSize, + candidateBlueState, + ref retainedBlueState); + + CopyCandidate( + candidateRedReconstruction, + candidateRedCoefficients, + redReconstruction, + chromaOrigin, + retainedRedCoefficients, + transformSize, + candidateRedState, + ref retainedRedState); + + for (int row = 0; row < height; row++) + { + colorIndexMap.DangerousGetRowSpan(row)[..width] + .CopyTo(retainedColorIndexMap[(row * width)..]); + } + + paletteInfo.PaletteSizes[1] = (byte)paletteSize; + paletteInfo.SetColors(Av1Plane.U, bluePaletteColors); + paletteInfo.SetColors(Av1Plane.V, redPaletteColors); + bestCost = candidateCost; + paletteSelected = true; + } + } + + if (paletteSelected) + { + for (int row = 0; row < height; row++) + { + retainedColorIndexMap.Slice(row * width, width) + .CopyTo(colorIndexMap.DangerousGetRowSpan(row)); + } + } + + return paletteSelected; + } + } +} diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs index 949ab13b39..5cfa4acbde 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs @@ -255,6 +255,7 @@ internal static partial class Av1IntraSuperblockEncoder redCoefficients[this.codedAreaChroma..], ref blueState, ref redState, + ref paletteInfo, out int chromaAngleDelta, out byte chromaFromLumaIndex, out sbyte chromaFromLumaSigns); diff --git a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1IntraSuperblockEncoderTests.cs b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1IntraSuperblockEncoderTests.cs index b5fb8ff788..4b02892b9c 100644 --- a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1IntraSuperblockEncoderTests.cs +++ b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1IntraSuperblockEncoderTests.cs @@ -780,6 +780,139 @@ public class Av1IntraSuperblockEncoderTests initialSize: 256)); } + [Fact] + public void ProductionTileSelectsExactPairedChromaPalette() + { + AssertProductionTileSelectsExactPairedChromaPalette(useLumaPalette: false); + AssertProductionTileSelectsExactPairedChromaPalette(useLumaPalette: true); + } + + private static void AssertProductionTileSelectsExactPairedChromaPalette(bool useLumaPalette) + { + const int Width = 8; + const int Height = 8; + const int QIndex = 37; + ObuColorConfig colorConfig = new() + { + IsMonochrome = false, + SubSamplingX = false, + SubSamplingY = false, + BitDepth = Av1BitDepth.EightBit + }; + + using Av1EncoderFrameBuffer source = new( + Configuration.Default, + Width, + Height, + 8, + Av1ColorFormat.Yuv444, + 0, + 0); + + using Av1EncoderFrameBuffer reconstruction = new( + Configuration.Default, + Width, + Height, + 8, + Av1ColorFormat.Yuv444, + 0, + 0); + + Buffer2DRegion lumaSource = source.Frame.CodedView.GetPlane(Av1Plane.Y); + Buffer2DRegion blueSource = source.Frame.CodedView.GetPlane(Av1Plane.U); + Buffer2DRegion redSource = source.Frame.CodedView.GetPlane(Av1Plane.V); + for (int row = 0; row < Height; row++) + { + Span lumaRow = lumaSource.DangerousGetRowSpan(row); + if (useLumaPalette) + { + for (int column = 0; column < Width; column++) + { + lumaRow[column] = column < Width / 2 ? (byte)64 : (byte)192; + } + } + else + { + lumaRow.Fill(128); + } + + blueSource.DangerousGetRowSpan(row).Fill(row < Height / 2 ? (byte)32 : (byte)224); + redSource.DangerousGetRowSpan(row).Fill(row < Height / 2 ? (byte)200 : (byte)40); + } + + ClearPlane(reconstruction.Luma); + ClearPlane(Assert.IsType>(reconstruction.ChromaBlue)); + ClearPlane(Assert.IsType>(reconstruction.ChromaRed)); + using Av1EncoderModeInfoBuffer modeInfo = new( + Configuration.Default, + Width, + Height, + disallow4x4AllFrames: true); + + Av1PictureControlSet pictureTemplate = CreatePicture( + modeInfo, + colorConfig, + use128x128Superblock: false, + QIndex); + + pictureTemplate.Parent.FrameHeader.AllowScreenContentTools = true; + pictureTemplate.Parent.FrameHeader.FrameSize.FrameWidth = Width; + pictureTemplate.Parent.FrameHeader.FrameSize.FrameHeight = Height; + using Av1EncoderPictureBuffer picture = new( + Configuration.Default, + pictureTemplate.Sequence.SequenceHeader, + pictureTemplate.Parent.FrameHeader, + Width, + Height); + + using Av1EncoderCoefficientBuffer coefficients = new( + Configuration.Default, + pictureTemplate.Sequence.SequenceHeader, + Width, + Height); + + using Av1EncoderSuperblockWorkspace superblockWorkspace = new(Configuration.Default); + using Av1EncoderBlockWorkspace blockWorkspace = new(Configuration.Default); + using Av1IntraTileWriter tileWriter = new( + Configuration.Default, + source.Frame, + reconstruction.Frame, + picture.Picture, + coefficients, + superblockWorkspace, + blockWorkspace, + initialSize: 256); + + ref Av1MacroBlockModeInfo mode = ref picture.Picture.GetMacroBlockModeInfo(default); + Assert.Equal(Av1ChromaPredictionMode.DC, mode.Block.UvMode); + Assert.Equal(useLumaPalette ? 2 : 0, superblockWorkspace.PaletteInfo.PaletteSizes[0]); + Assert.Equal(2, superblockWorkspace.PaletteInfo.PaletteSizes[1]); + Assert.Equal([32, 224], superblockWorkspace.PaletteInfo.GetColors(Av1Plane.U).ToArray()); + Assert.Equal([200, 40], superblockWorkspace.PaletteInfo.GetColors(Av1Plane.V).ToArray()); + Assert.Equal((ushort)0, coefficients.GetTransformBlockSpan(0, Av1Plane.U)[0].EndOfBlock); + Assert.Equal((ushort)0, coefficients.GetTransformBlockSpan(0, Av1Plane.V)[0].EndOfBlock); + + Buffer2DRegion colorIndexMap = superblockWorkspace + .GetPaletteMaps() + .GetMap(Av1PlaneType.Uv, Width, Height); + + Buffer2DRegion blueReconstruction = reconstruction.Frame.CodedView.GetPlane(Av1Plane.U); + Buffer2DRegion redReconstruction = reconstruction.Frame.CodedView.GetPlane(Av1Plane.V); + for (int row = 0; row < Height; row++) + { + byte expectedIndex = (byte)(row < Height / 2 ? 0 : 1); + foreach (byte index in colorIndexMap.DangerousGetRowSpan(row)) + { + Assert.Equal(expectedIndex, index); + } + + Assert.True(blueSource.DangerousGetRowSpan(row).SequenceEqual(blueReconstruction.DangerousGetRowSpan(row))); + Assert.True(redSource.DangerousGetRowSpan(row).SequenceEqual(redReconstruction.DangerousGetRowSpan(row))); + } + + Assert.NotEqual(0, tileWriter.GetTileData(0).Length); + } + [Theory] [InlineData((int)Av1PredictionMode.Vertical, 0)] [InlineData((int)Av1PredictionMode.Horizontal, 0)]