From cab9bc47cf366ebe6f2c07861b0314ba77d14ea7 Mon Sep 17 00:00:00 2001 From: James Jackson-South Date: Thu, 3 Sep 2026 21:30:01 +1000 Subject: [PATCH] Complete exhaustive luma transform search --- HEIF_IMPLEMENTATION_PLAN.md | 8 +- .../Av1EncoderModeDecisionWorkspace.cs | 8 +- ...traSuperblockEncoder.ChromaModeDecision.cs | 10 + ...rblockEncoder.ChromaPaletteModeDecision.cs | 10 + ...blockEncoder.IntraBlockCopyModeDecision.cs | 46 ++- .../Av1IntraSuperblockEncoder.ModeDecision.cs | 311 +++++++++++++----- ...raSuperblockEncoder.PaletteModeDecision.cs | 14 +- .../Formats/Heif/Av1/Av1EncoderFrameTests.cs | 3 + 8 files changed, 322 insertions(+), 88 deletions(-) diff --git a/HEIF_IMPLEMENTATION_PLAN.md b/HEIF_IMPLEMENTATION_PLAN.md index 2d216272ac..366181afa5 100644 --- a/HEIF_IMPLEMENTATION_PLAN.md +++ b/HEIF_IMPLEMENTATION_PLAN.md @@ -829,8 +829,8 @@ Encoder verification contract: - [~] Implement intra mode search, palette, filter intra, chroma-from-luma, and intra-block copy decisions. Live luma search now covers all 13 zero-angle base modes and all six nonzero adjustments for each of the eight directional modes. Joint spatial chroma search covers the same 61 candidates, combines both chroma planes in one rate-distortion decision, and preserves the winning shared angle adjustment. Chroma-from-luma now searches the complete signed alpha alphabet from reconstructed luma and retains its joint U/V syntax. Filter-intra now searches all five predictors after ordinary luma modes. Palette entropy, retained state, production syntax, exhaustive luma and paired chroma palette selection, adaptive screen-content activation, and joint intra-block-copy mode selection are complete. - [ ] Implement inter mode search for bounded sequences, including reference selection and the decoder-supported inter tools. - [~] Current-libaom `av1_quantize_fp_no_qmatrix` arithmetic is implemented as a closed generic forward-quantizer family with Vector512, Vector256, Vector128, and scalar paths, raster-order output, coded 64-point coefficient limits, and scan-order EOB selection. Transform search, coefficient optimization, and lossless behavior remain. -- [~] Implement real rate-distortion selection and make quality and effort change work, size, and output quality. The complete luma and joint chroma candidate sets, including chroma-from-luma, filter-intra, palette, and intra-block copy, now perform live rate-distortion selection; quality mapping, effort tiers above the current uniform-transform search ceiling, and the remaining searches are not implemented. -- [~] Frame effort now progressively expands the available current search: zero is DC-only, one adds every zero-angle spatial mode, two adds every legal directional adjustment, three adds transform refinement, four adds filter-intra and chroma-from-luma, and five adds adaptive palette and intra-block-copy analysis. Lower tiers do not signal unavailable sequence or frame tools, and tiers below five skip the whole-frame screen-content scan. Effort six enables `TX_MODE_SELECT` and compares the winning ordinary spatial or filter-intra luma mode as one 8x8 transform against four raster-ordered 4x4 transforms. Each luma palette candidate performs that same transform-size comparison before competing with other palette sizes. Every 4x4 transform searches all legal transform types with live coefficient contexts and reconstructed intra references. The search reuses the aligned block workspace, preserves only improving candidates, and performs no per-block or per-transform rent. Non-skipped intra-block copy writes and costs the current-libaom unsplit variable-transform root; skipped intra-block copy emits no transform-partition symbol. Values seven through ten currently share the effort-six ceiling. Broader joint mode/transform refinement, partition search, and pruning remain. Decoder-visible production cases verify emitted flags, mode restrictions, real 4x4 selection, filter-intra, palette, intra-block-copy syntax, and successful decode. The latest incremental net11 Release build reports three warnings and zero errors, and the complete non-HEVC HEIF/AV1 namespace passes 9,298 of 9,298 through direct foreground VSTest. Current-main `aomdec` at `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727` accepts the generated 8-bit and 12-bit palette transform-size payloads and the effort-six intra-block-copy payload. Roslynk reports zero compiler errors. +- [~] Implement real rate-distortion selection and make quality and effort change work, size, and output quality. The complete luma and joint chroma candidate sets, including chroma-from-luma, filter-intra, palette, and intra-block copy, now perform live rate-distortion selection. Public quality mapping and effort tiers through exhaustive uniform luma mode/transform search are implemented; partition search, effort-dependent pruning, and the remaining sequence searches remain. +- [~] Frame effort now progressively expands the available current search: zero is DC-only, one adds every zero-angle spatial mode, two adds every legal directional adjustment, three refines the preliminary luma winner's transform type, four adds filter-intra and chroma-from-luma, and five adds adaptive palette and intra-block-copy analysis. Lower tiers do not signal unavailable sequence or frame tools, and tiers below five skip the whole-frame screen-content scan. Effort six enables `TX_MODE_SELECT` and compares the winning ordinary spatial or filter-intra luma mode as one 8x8 transform against four raster-ordered 4x4 transforms; each luma palette candidate owns that size comparison from effort six onward. Effort seven searches every legal 8x8 transform type inside every ordinary spatial candidate rather than refining only the preliminary winner. Effort eight also performs the 8x8-versus-four-4x4 comparison inside every ordinary spatial and filter-intra candidate, matching current libaom's per-candidate uniform-transform ownership. Values nine and ten currently share the effort-eight ceiling. Prediction and residual construction run once per mode and are reused across its legal transform types. Every 4x4 transform searches all legal types with live coefficient contexts and reconstructed intra references. The search reuses the aligned block workspace, preserves only global improvements, and performs no per-block or per-transform rent. Non-skipped intra-block copy writes and costs the current-libaom unsplit variable-transform root; skipped intra-block copy emits no transform-partition symbol. Partition search and effort-dependent model/transform pruning remain. Decoder-visible production cases now execute effort zero through eight and ten, inspect the emitted restrictions and frame state, and decode the produced streams. The complete non-HEVC HEIF/AV1 namespace passes 9,301 of 9,301 through one foreground net11 VSTest run. Current-main `aomdec` at `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727` accepts the generated effort-eight and effort-ten payloads as well as the existing palette and intra-block-copy payloads. The net11 Release build and Roslynk compiler pass report zero errors. - [~] Encoder rate accounting converts the entropy writer's live inverse cumulative distributions into current-libaom fixed-point symbol costs without allocating or duplicating probability state. Read-only luma-mode, directional-delta, filter-intra, chroma-mode, block-skip, transform-size, transform-block-skip, and complete transform-coefficient queries share the exact distributions mutated by the subsequent entropy write. Complete coefficient costing follows current libaom's optimized shape: it returns immediately for an empty transform, uses the EOB-specific base-range context, fuses magnitude, sign, base-range, and Golomb accounting into one reverse traversal, and combines repeated full base-range chunks instead of replaying each emitted symbol. Tile-lifetime level and context scratch is reused, the one-coefficient path neither clears nor initializes the forward-neighbor level map, and steady-state queries allocate nothing. Transform-size writing and costing share one subdivision-depth calculation, while shared closed symbol operations keep the writer and cost mappings for transform skip, transform type, and EOB syntax identical without forcing the estimator through the writer's slower two-pass coefficient traversal. The current-libaom fixed-point RD combiner preserves 64-bit distortion and rounds the weighted 1/512-bit rate at the required boundary. Its key-frame multiplier follows libaom's squared DC-quantizer formula and exact 10/12-bit normalization. Live final-block selection evaluates all 61 legal 8x8 luma candidates: the 13 zero-angle base modes in current-libaom order, followed by six nonzero adjustments for each directional mode. Joint chroma selection evaluates the equivalent 61 spatial candidates, combines U and V distortion plus coefficient rate, and charges one live chroma-mode and shared-angle symbol over the actual subsampled 4x4, 4x8, or 8x8 geometry. Chroma-from-luma subsamples the reconstructed luma block once into fixed-stride Q3 stack scratch, subtracts the rounded mean, evaluates all 33 signed alpha values independently for each plane with complete transform RD, and combines the cached plane results across all 1,088 valid joint pairs with one live sign cost and the conditional U/V magnitude costs. This is the allocation-free equivalent of current libaom's exhaustive 33-value path: it requires 66 evaluation transforms rather than transforming every joint pair, preserves DC-before-CfL-before-spatial tie order, and fixes the implicit chroma transform to DCT-DCT. Filter-intra follows ordinary luma candidates, searches all five predictors in syntax order, and evaluates every legal transform while reusing one prepared prediction and source residual per filter mode. Every candidate includes its live mode, angle, filter mode, alpha, and coefficient rate plus normalized pixel-domain distortion. Each prepared reference edge retains the common-corner prefix and twice the transform dimension required by directional prediction. A shared encoder/decoder availability calculation selects reconstructed top-right and bottom-left extensions according to tile, frame, superblock, and block reconstruction order; unavailable extensions repeat the nearest coded endpoint. Missing top or left edges retain current libaom's perpendicular-sample and bit-depth-midpoint rules. Directional prediction applies the AV1 three-degree adjustment step and reuses transform workspace for zone-three transposition before the transform overwrites it, keeping candidate evaluation allocation-free. The winning luma and chroma signed adjustments are retained in the packed final-block state consumed by the tile writer. The tile writer invokes these reusable workspace-backed selectors after mapping current neighbors and immediately before writing each block, so later decisions see reconstructed samples, coefficient contexts, and CDF updates from every preceding block. Block skip is read only after the callback has combined every coded plane. Luma and chroma candidate scratch is partitioned from the encoder's single aligned reusable block workspace; transform-size search uses that owner for four retained 4x4 transform states, local coefficient contexts, and the compact trial reconstruction needed to preserve the best result. No candidate path rents a buffer per block or per transform. Only a newly winning candidate is copied into retained frame storage. Production fixtures force every luma base predictor, both extreme adjustments in all three directional zones, available top-right and bottom-left extensions, high-bit-depth adjustment propagation, exact signed luma and chroma angle-rate terms, joint U/V decisions, packed chroma state, and 4:2:0, 4:2:2, and 4:4:4 transform geometry. The CfL fixtures derive target chroma from a pilot production encode's actual reconstructed luma through an independent scalar Q3 oracle and prove exact positive/negative alpha syntax plus zero-residual DCT-DCT reconstruction for all three subsampling geometries at 8, 10, and 12 bits. The stable fixed-DC traversal comparison uses neutral samples for which both the baseline and live search are contractually DC and skipped, instead of relying on textured content to happen to select the baseline mode. Luma palette selection now evaluates dominant-color and one-dimensional K-means candidates for every legal size, snaps near-cache colors with the reference threshold and tie order, removes duplicate snapped colors, extends boundary maps from active samples, and performs complete transform rate-distortion search. Ordinary DC and filter-intra candidates pay the palette-disabled symbol whenever screen-content syntax is enabled. The exact net11 Release rebuild reports 1,992 test-project warnings and zero errors, all 58 intra-superblock cases pass, all 8,935 AVIF cases pass, and all 230 HEIF cases pass. Remaining mode decision work includes transform-size coverage for filter-intra and palette, broader joint mode/transform refinement, partition search, and effort-dependent pruning. Non-empty intra blocks deliberately remain non-skipped, matching current libaom; later inter mode selection owns its distinct skip-transform RD decision. - [~] The tile writer now publishes one packed coefficient context per covered 4x4 edge unit and derives luma/chroma skip plus DC-sign contexts from the complete transform edges using current-libaom units. Partition, transform, and coefficient neighbor state retains only the above and left context regions used by current libaom; the unused third top-left region, its granularity state, and its unused sentinel are removed. One picture owner now packs segmentation plus every tile's partition, luma, chroma, and transform edges into one clean byte allocation with typed non-owning views; together with the separately typed packed mode-information owner, the complete picture state uses two allocator rents rather than seven. Exact aligned lengths, clean initialization, and balanced exactly-once returns are covered in Release. Multi-tile payload ownership and verified CDF update behavior remain. - [~] Encoder mode information now uses a frame-owned integer alias grid over a packed 8-byte value allocation, matching current libaom's `mi_grid_base` and `mi_alloc` relationship without a managed object or reference per 4x4 entry. The visible dimensions are aligned to eight luma samples, the grid stride and allocated row count are aligned to 32 mode-information units, and optional 8x8 allocation granularity reduces the value store in both dimensions exactly as current libaom does. One clean ImageSharp byte owner contains both independently typed regions, reducing libaom's two allocation lifetimes to one without a copy. At 4K, the 4x4 layout occupies about 6.0 MiB in total; the 8x8 layout occupies about 3.0 MiB. Exact geometry, clean allocation, typed lengths, aligned mapping, untouched row padding, and exactly-once return pass 4 of 4 direct net11 VSTest cases in Release. Every coded 4x4 cell covered by square, rectangular, or clipped edge blocks maps to its owning allocation entry before context-dependent symbols are written. Packed syntax, relative neighbor lookup, full block mapping, writer traversal, entropy, and OBU coverage pass 1,947 of 1,947 direct net11 VSTest cases in Release; complete mode decision still remains. @@ -868,7 +868,9 @@ Encoder verification contract: - [x] Operation-wide allocation tracking now exercises a real 64x64 12-bit 4:4:4 frame through packed-pixel conversion, both native frame owners, picture and coefficient state, reusable block workspaces, entropy coding, OBU framing, and a non-seekable destination. It proves exactly one 60 KiB tile-output reservation from current libaom's all-intra 2.5x rule and balanced exactly-once returns for every tracked allocation before the operation completes. The focused ownership case passes 1 of 1 and the complete HEIF/AV1 namespace passes 8,863 of 8,863 direct net11 VSTest cases with zero failures or skips. - [x] HEIF box offsets are now counted from the start of the encoded file instead of reading `Stream.Position`. This preserves ISO BMFF file-relative `iloc` offsets when the destination begins at a nonzero position and permits non-seekable output. Decoder item extents and image-sequence chunk offsets now resolve from that same file origin rather than the backing stream origin. Real legacy-JPEG HEIF round trips cover non-seekable output and a prefixed destination, while current-position AV1 decode covers both a still item and a five-frame sequence. All 96 encoder/decoder cases and all 38 sequence-parser cases pass direct net11 Release VSTest; the Release build remains at the established 1,005-warning baseline with zero errors. -- [~] Effort-six uniform luma transform selection now compares ordinary spatial and filter-intra winners, plus every individual luma palette candidate, as one 8x8 transform against four raster-ordered 4x4 transforms. Searching transform size inside each palette candidate matches current libaom's per-candidate uniform-transform decision and prevents an 8x8-only preliminary result from discarding the palette that is optimal with 4x4 residuals. Each 4x4 transform searches every legal type with coefficient contexts derived from retained transform edges and preceding trial blocks, while reconstructed top-right and bottom-left references follow production coding order. Palette prediction uses non-owning subregions of the retained color map, and filter-intra rebuilds each recursive prediction from reconstructed edges. The strided transform operator writes directly into the 8x8 reconstruction mosaic, and the existing aligned block-workspace owner retains prediction, residual, coefficients, contexts, compact reconstruction, and four final states; no allocator rent, managed array, best-candidate re-transform, or full-block intermediate copy was added. The source comparison also exposed that encoder transform edges began at zero while libaom and the ImageSharp decoder initialize unavailable edges to the largest transform. The packed encoder edge storage now initializes to 64, so variable-transform partition contexts agree before a coded neighbor publishes its size. Variable transform syntax is also gated to blocks larger than 4x4, matching libaom's `block_signals_txsize`. The focused Release verification passes 5 of 5 cases, including 8-bit and 12-bit palette split selection, clipped palette blocks, intra-block copy, and transform-size edge contexts. The complete non-HEVC HEIF/AV1 namespace passes 9,298 of 9,298 cases with zero failures or skips. Current-main `aomdec` at `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727` accepts both generated palette transform-size streams and the effort-six intra-block-copy stream. Broader joint mode/transform refinement, partition search, and effort-dependent pruning remain. +- [~] Current-libaom source comparison now drives uniform luma transform ownership at each effort boundary. Effort six retains the cheaper winner-only size decision for ordinary spatial and filter-intra modes, while every palette candidate already owns its size decision. Effort seven evaluates every legal 8x8 transform type for every ordinary spatial candidate. Effort eight and above make transform size part of every ordinary spatial and filter-intra candidate's rate-distortion result, so an 8x8-only preliminary comparison cannot discard the mode that wins with four 4x4 transforms. Prediction and subtraction are prepared once per mode and reused across transform types, matching the reference separation between prediction and transform search. Each 4x4 transform searches every legal type with coefficient contexts derived from retained transform edges and preceding trial blocks, while reconstructed top-right and bottom-left references follow production coding order. Palette prediction uses non-owning subregions of the retained color map, and filter-intra rebuilds each recursive prediction from reconstructed edges. The existing aligned block-workspace owner retains prediction, residual, coefficients, contexts, compact reconstruction, and four final states; no allocator rent, managed array, best-candidate re-transform, or full-block intermediate copy was added. Dense decision points now document scratch lifetime, enumeration tie order, global-winner publication, raster reconstruction dependencies, and the deliberate lower-effort shortcut. The packed encoder transform edges initialize to 64, matching libaom and the ImageSharp decoder before a coded neighbor publishes its size, and variable transform syntax remains gated to blocks larger than 4x4. The focused Release verification passes 13 of 13 cases across efforts zero through eight and ten, palette split selection, and transform-size selection. The complete non-HEVC HEIF/AV1 namespace passes 9,301 of 9,301 cases with zero failures or skips. Current-main `aomdec` at `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727` accepts the generated effort-eight and effort-ten streams. Partition search and effort-dependent pruning remain. + +- [x] Intra-block-copy transform search now prepares motion compensation and subtraction once per plane, alternates the existing candidate and selected work buffers whenever a transform improves, and performs at most one final normalization copy into the caller-owned selected span. This matches current libaom's pointer-swap ownership without adding an allocation or a third reconstruction buffer. Inline documentation now records the scratch lifetime, strict transform tie order, skip-rate replacement, unsplit transform-root syntax, joint-plane winner retention, and final publication boundary. The focused Release encoder and intra-block-copy set passes 17 of 17 cases, the complete non-HEVC HEIF/AV1 namespace passes 9,301 of 9,301 cases with zero failures or skips, and current-main `aomdec` accepts the regenerated effort-five and effort-six intra-block-copy streams. ### 7. Write complete AVIF output diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs index 8ec5f48f4f..f93708151d 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs @@ -64,15 +64,15 @@ internal readonly ref struct Av1EncoderModeDecisionWorkspace public Av1EncoderModeDecisionWorkspace(Span storage) => this.storage = storage; /// - /// Gets the temporary prediction span used by filter-intra mode search. + /// Gets the temporary prediction span shared by mutually exclusive mode searches. /// - public Span FilterPrediction + public Span Prediction => MemoryMarshal.Cast(this.storage[TransientStorageOffset..])[..MaximumSampleCount]; /// - /// Gets the temporary residual span used by filter-intra mode search. + /// Gets the temporary residual span shared by mutually exclusive mode searches. /// - public Span FilterResidual + public Span Residual => MemoryMarshal.Cast( this.storage.Slice( TransientStorageOffset + (MaximumSampleCount * sizeof(ushort) / sizeof(int)), diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs index df02652215..6ab322a888 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs @@ -200,6 +200,9 @@ internal static partial class Av1IntraSuperblockEncoder Av1EncoderTransformBlockState candidateBlueState = default; Av1EncoderTransformBlockState candidateRedState = default; + + // A chroma mode and angle are shared by U and V, so neither plane can replace the + // retained result independently. Their complete rate and distortion compete jointly. long candidateCost = this.GetChromaCandidateCost( writer, modeInfo, @@ -395,6 +398,8 @@ internal static partial class Av1IntraSuperblockEncoder if (chromaFromLumaSelected) { + // The alpha tables retain only rate and distortion. Regenerate the two selected planes + // once here instead of copying reconstruction and coefficient blocks for all 66 trials. Av1EncoderTransformBlockState candidateBlueState = default; _ = this.GetChromaFromLumaPlaneCost( writer, @@ -561,6 +566,9 @@ internal static partial class Av1IntraSuperblockEncoder { const Av1BlockSize BlockSize = Av1BlockSize.Block8x8; Av1PredictionMode predictionMode = chromaMode.ToLumaMode(); + + // Intra chroma derives one transform type from the shared UV prediction mode. The type is not + // signaled independently for either chroma plane, so U and V must use the same legal fallback. Av1TransformType transformType = Av1SymbolContextHelper.GetDefaultIntraTransformType( predictionMode, transformSize, @@ -608,6 +616,8 @@ internal static partial class Av1IntraSuperblockEncoder this.bitDepth, ref candidateRedState); + // The mode and angle are written once for the UV pair; coefficient syntax remains independent + // because each plane has its own EOB, scan values, and neighboring coefficient context. int rate = Av1TileWriter.GetChromaModeCost( writer, this.picture.Parent.FrameHeader, diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaPaletteModeDecision.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaPaletteModeDecision.cs index f67f950310..1afa5f6140 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaPaletteModeDecision.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaPaletteModeDecision.cs @@ -150,6 +150,8 @@ internal static partial class Av1IntraSuperblockEncoder workspace.GetAlternateCentroids(1), workspace.AlternateIndices); + // Only U participates in the neighbor-color cache. Snap bounded U deltas before sorting, + // while preserving each U/V centroid as one paired palette entry. for (int colorIndex = 0; colorIndex < paletteSize && !colorCache.IsEmpty; colorIndex++) { int minimumDifference = Math.Abs(candidateBlueCentroids[colorIndex] - colorCache[0]); @@ -221,6 +223,8 @@ internal static partial class Av1IntraSuperblockEncoder redPaletteColors[colorIndex] = (ushort)candidateRedCentroids[colorIndex]; } + // U and V share one color-index map but reconstruct through their own palette values and + // residuals. Both preparations remain valid until the next palette-size candidate. TOperator.PreparePalette( blueSource, chromaOrigin, @@ -239,6 +243,8 @@ internal static partial class Av1IntraSuperblockEncoder redResidual[..sampleCount], transformSize); + // Intra chroma derives one transform type from the shared UV mode. Palette uses UV DC, so + // both planes use DCT while retaining independent coefficient contexts and end positions. Av1EncoderTransformBlockState candidateBlueState = default; long distortion = TOperator.EncodePredictionCandidate( this.blockWorkspace, @@ -314,6 +320,8 @@ internal static partial class Av1IntraSuperblockEncoder Av1FilterIntraMode.AllFilterIntraModes, usesInterTransformSet: false); + // Mode, palette, and color-map syntax is shared by the pair; coefficient syntax and + // distortion remain per plane before the joint chroma rate-distortion comparison. rate += writer.GetCoefficientCost( transformSize, Av1TransformType.DctDct, @@ -329,6 +337,8 @@ internal static partial class Av1IntraSuperblockEncoder long candidateCost = Av1RateDistortion.GetCost(this.rateMultiplier, rate, distortion); if (candidateCost < bestCost) { + // Every following palette size overwrites the shared maps and candidate spans, so a + // global improvement must retain reconstruction, coefficients, colors, and indices together. Buffer2DRegion blueReconstruction = this.reconstruction.GetPlane(Av1Plane.U); Buffer2DRegion redReconstruction = this.reconstruction.GetPlane(Av1Plane.V); CopyCandidate( diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.IntraBlockCopyModeDecision.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.IntraBlockCopyModeDecision.cs index 6ba0014779..5c244b07a0 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.IntraBlockCopyModeDecision.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.IntraBlockCopyModeDecision.cs @@ -102,6 +102,8 @@ internal static partial class Av1IntraSuperblockEncoder return; } + // Mode-decision costs already contain the selected coefficient syntax but not the block's skip + // or IBC choice. When regular intra skips, replace its empty-coefficient rate with skip syntax. int skipContext = Av1TileWriter.GetSkipContext(macroBlock); int regularRateAdjustment = writer.GetUseIntraBlockCopyCost(false) + writer.GetSkipCost(modeInfo.Block.Skip, skipContext); @@ -128,6 +130,8 @@ internal static partial class Av1IntraSuperblockEncoder BlockSize, LumaTransformSize); + // A coded IBC residual uses the unsplit transform root at this fixed block size. A skipped block + // omits both the transform-partition bit and coefficient syntax, so this rate is added only below. int transformPartitionRate = 0; if (this.picture.Parent.FrameHeader.TransformMode == Av1TransformMode.Select) { @@ -176,6 +180,8 @@ internal static partial class Av1IntraSuperblockEncoder chromaTransformSize); } + // Per-vector plane results reuse candidate scratch. Separate selected spans retain only a new + // global winner, allowing the complete search to finish before committed reconstruction changes. for (int candidateIndex = 0; candidateIndex < uniqueCandidateCount; candidateIndex++) { Av1MotionVector candidate = candidates[candidateIndex]; @@ -291,6 +297,9 @@ internal static partial class Av1IntraSuperblockEncoder long candidateDistortion = lumaDistortion + blueDistortion + redDistortion; long candidateCost = Av1RateDistortion.GetCost(this.rateMultiplier, candidateRate, candidateDistortion); bool candidateSkip = false; + + // The skip alternative is available only when every coded plane has an empty transform. Its + // distortion comes from prediction alone and its rate excludes the transform tree and coefficients. if (hasEmptyLuma && hasEmptyBlue && hasEmptyRed) { int skipRate = writer.GetUseIntraBlockCopyCost(true) + @@ -363,6 +372,8 @@ internal static partial class Av1IntraSuperblockEncoder return; } + // Only the winning vector is now visible to later coding blocks. This single publication keeps + // rejected motion vectors from contaminating intra references or entropy contexts. Span retainedLumaCoefficients = this.coefficientBuffer.GetPlaneSpan(this.superblock.Index, Av1Plane.Y); Span retainedLumaTransformBlocks = this.coefficientBuffer.GetTransformBlockSpan(this.superblock.Index, Av1Plane.Y); @@ -473,6 +484,8 @@ internal static partial class Av1IntraSuperblockEncoder residual[..sampleCount], transformSize); + // Motion compensation and subtraction do not depend on transform type. Keep them outside the + // transform loop so exhaustive luma search traverses the source and reference blocks only once. Av1TransformSetType transformSetType = Av1SymbolContextHelper.GetExtendedTransformSetType( transformSize, isInter: true, @@ -492,6 +505,15 @@ internal static partial class Av1IntraSuperblockEncoder hasEmptyTransform = false; emptyState = default; emptyDistortion = 0; + + // The candidate and best spans alternate ownership whenever a transform improves the result. + // This mirrors the reference's buffer-pointer swap and replaces a copy on every improvement + // with at most one normalization copy after the transform search. + Span candidateReconstruction = transformReconstruction[..sampleCount]; + Span candidateCoefficients = transformCoefficients[..sampleCount]; + Span bestReconstruction = selectedReconstruction[..sampleCount]; + Span bestCoefficients = selectedCoefficients[..sampleCount]; + bool bestUsesSelectedStorage = true; for (Av1TransformType transformType = firstTransformType; transformType < transformTypeLimit; transformType++) @@ -508,9 +530,9 @@ internal static partial class Av1IntraSuperblockEncoder planeOrigin, prediction[..sampleCount], residual[..sampleCount], - transformReconstruction[..sampleCount], + candidateReconstruction, transformSize.GetWidth(), - transformCoefficients[..sampleCount], + candidateCoefficients, transformSize, transformType, plane, @@ -524,7 +546,7 @@ internal static partial class Av1IntraSuperblockEncoder transformSize, transformType, Av1PredictionMode.DC, - transformCoefficients[..sampleCount], + candidateCoefficients, componentType, blockContext, candidateState.EndOfBlock, @@ -539,8 +561,14 @@ internal static partial class Av1IntraSuperblockEncoder if (candidateCost < bestCost) { - transformReconstruction[..sampleCount].CopyTo(selectedReconstruction); - transformCoefficients[..sampleCount].CopyTo(selectedCoefficients); + Span previousBestReconstruction = bestReconstruction; + bestReconstruction = candidateReconstruction; + candidateReconstruction = previousBestReconstruction; + + Span previousBestCoefficients = bestCoefficients; + bestCoefficients = candidateCoefficients; + candidateCoefficients = previousBestCoefficients; + bestUsesSelectedStorage = !bestUsesSelectedStorage; bestCost = candidateCost; selectedState = candidateState; selectedRate = candidateRate; @@ -555,6 +583,14 @@ internal static partial class Av1IntraSuperblockEncoder emptyDistortion = candidateDistortion; } } + + // Callers retain the designated selected spans after this scratch workspace is reused by the + // next plane or motion vector, so normalize only when the final best result occupies scratch. + if (!bestUsesSelectedStorage) + { + bestReconstruction.CopyTo(selectedReconstruction); + bestCoefficients.CopyTo(selectedCoefficients); + } } } } diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs index ca9d5b0629..0847a7350a 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs @@ -207,6 +207,8 @@ internal static partial class Av1IntraSuperblockEncoder block.FilterIntraMode = filterIntraMode; modeInfo.Block.TransformSize = lumaTransformSize; + // One 8x8 coding block retains either one 8x8 transform or four 4x4 transforms. The block-level + // skip decision is legal only when every transform selected by mode decision has an empty EOB. int lumaTransformBlockCount = LumaTransformSize.GetSize2d() / lumaTransformSize.GetSize2d(); bool lumaTransformEmpty = true; for (int transformIndex = 0; transformIndex < lumaTransformBlockCount; transformIndex++) @@ -295,6 +297,8 @@ internal static partial class Av1IntraSuperblockEncoder block.PredictionUnit.ChromaFromLumaIndex = chromaFromLumaIndex; block.PredictionUnit.ChromaFromLumaSigns = chromaFromLumaSigns; + // Skip suppresses coefficient syntax for the entire coding block, not one plane independently. + // Preserve normal coefficient coding when any selected luma or chroma transform is nonempty. bool allTransformsEmpty = lumaTransformEmpty && blueState.EndOfBlock == 0 && redState.EndOfBlock == 0; int regularEmptyTransformRate = 0; if (allTransformsEmpty) @@ -590,10 +594,13 @@ internal static partial class Av1IntraSuperblockEncoder Span candidateReconstruction = workspace.GetCandidateReconstruction(0); Span candidateCoefficients = workspace.GetCandidateCoefficients(0); + Span prediction = workspace.Prediction; + Span residual = workspace.Residual; long bestCost = long.MaxValue; Av1PredictionMode bestMode = Av1PredictionMode.DC; selectedAngleDelta = 0; selectedFilterIntraMode = Av1FilterIntraMode.AllFilterIntraModes; + selectedTransformSize = TransformSize; int baseModeCount = LumaModeSearchOrder.Length; int deltaCount = AngleDeltaSearchOrder.Length; int directionalModeCount = (int)Av1PredictionMode.Directional67Degrees - (int)Av1PredictionMode.Vertical + 1; @@ -611,6 +618,11 @@ internal static partial class Av1IntraSuperblockEncoder TransformSize, useReducedTransformSet); + // Transform type and transform size are separate search axes. Splitting their effort thresholds + // gives callers a useful intermediate tier without changing the fast default path. + bool searchEveryTransformType = this.effort >= 7; + bool searchEveryTransformSize = this.effort >= 8; + // Zero-angle modes precede groups of six nonzero adjustments for each directional mode. // A single index preserves that tie-breaking order without duplicating candidate evaluation. for (int candidateIndex = 0; candidateIndex < candidateCount; candidateIndex++) @@ -629,102 +641,196 @@ internal static partial class Av1IntraSuperblockEncoder angleDelta = AngleDeltaSearchOrder[adjustedIndex % deltaCount]; } - Av1TransformType defaultTransformType = Av1SymbolContextHelper.GetDefaultIntraTransformType( - mode, - TransformSize, - useReducedTransformSet); - - Av1EncoderTransformBlockState candidateState = default; - long candidateCost = this.GetLumaCandidateCost( - writer, - macroBlock, + // Prediction and subtraction do not depend on transform type. Preparing them once keeps + // exhaustive transform search from repeating the same pixel traversal for every candidate. + TOperator.PrepareIntra( + this.blockWorkspace, sourcePlane, blockOrigin, + prediction, above, left, hasLeft, hasAbove, mode, angleDelta, - defaultTransformType, - blockContext, - paletteDisabledCost, - largestTransformRate, - candidateReconstruction, - candidateCoefficients, - ref candidateState); + residual, + TransformSize, + this.bitDepth); + + // Transform types are visited in AV1 enumeration order. A strict cost comparison below keeps + // the first legal type on ties, while lower efforts visit only the mode-derived default. + Av1TransformType firstTransformType = searchEveryTransformType + ? Av1TransformType.DctDct + : Av1SymbolContextHelper.GetDefaultIntraTransformType( + mode, + TransformSize, + useReducedTransformSet); - if (candidateCost < bestCost) + Av1TransformType transformTypeLimit = searchEveryTransformType + ? Av1TransformType.AllTransformTypes + : (Av1TransformType)((int)firstTransformType + 1); + + for (Av1TransformType transformType = firstTransformType; + transformType < transformTypeLimit; + transformType++) { - CopyCandidate( + if (!transformType.IsExtendedSetUsed(transformSetType)) + { + continue; + } + + Av1EncoderTransformBlockState candidateState = default; + long candidateCost = this.GetLumaCandidateCost( + writer, + macroBlock, + sourcePlane, + blockOrigin, + prediction, + residual, + mode, + angleDelta, + transformType, + blockContext, + paletteDisabledCost, + largestTransformRate, candidateReconstruction, candidateCoefficients, + ref candidateState); + + if (candidateCost < bestCost) + { + // The shared candidate spans are overwritten by the next transform. Copy only a + // global improvement into final block storage so no per-mode retained buffer is needed. + CopyCandidate( + candidateReconstruction, + candidateCoefficients, + reconstructionPlane, + blockOrigin, + retainedCoefficients, + TransformSize, + candidateState, + ref retainedStates[0]); + + bestCost = candidateCost; + bestMode = mode; + selectedAngleDelta = angleDelta; + selectedTransformSize = TransformSize; + } + } + + // At exhaustive effort, transform size belongs to this mode's RD result. Evaluate it + // before advancing so an 8x8-only preliminary result cannot discard a better split mode. + if (searchEveryTransformSize && + this.picture.Parent.FrameHeader.TransformMode == Av1TransformMode.Select) + { + long splitCost = this.GetSplitLumaCandidateCost( + writer, + macroBlock, + sourcePlane, reconstructionPlane, blockOrigin, - retainedCoefficients, - TransformSize, - candidateState, - ref retainedStates[0]); + tileIndex, + mode, + angleDelta, + Av1FilterIntraMode.AllFilterIntraModes, + 0, + ReadOnlySpan.Empty, + 0, + paletteDisabledCost, + transformSizeContext, + bestCost, + candidateReconstruction, + candidateCoefficients, + workspace.CandidateTransformBlocks); - bestCost = candidateCost; - bestMode = mode; - selectedAngleDelta = angleDelta; + if (splitCost < bestCost) + { + CopySplitCandidate( + candidateReconstruction, + candidateCoefficients, + workspace.CandidateTransformBlocks, + reconstructionPlane, + blockOrigin, + retainedCoefficients, + retainedStates); + + bestCost = splitCost; + bestMode = mode; + selectedAngleDelta = angleDelta; + selectedTransformSize = Av1TransformSize.Size4x4; + } } } - // Lower effort levels retain the mode-derived default transform. Higher levels refine only the winning - // mode across the legal transform set, preserving search order without evaluating every mode-transform pair. + // Midrange effort refines the preliminary mode only. Higher effort already searched every + // mode-transform pair above, so repeating the winning mode would add no candidates. long bestTransformCost = bestCost; - for (Av1TransformType transformType = Av1TransformType.DctDct; - this.effort >= 3 && transformType < Av1TransformType.AllTransformTypes; - transformType++) + if (this.effort >= 3 && !searchEveryTransformType) { - if (!transformType.IsExtendedSetUsed(transformSetType)) - { - continue; - } - - Av1EncoderTransformBlockState candidateState = default; - long candidateCost = this.GetLumaCandidateCost( - writer, - macroBlock, + // The shared spans now contain the last mode visited above, so rebuild the preliminary + // winner once before refining its transform types. + TOperator.PrepareIntra( + this.blockWorkspace, sourcePlane, blockOrigin, + prediction, above, left, hasLeft, hasAbove, bestMode, selectedAngleDelta, - transformType, - blockContext, - paletteDisabledCost, - largestTransformRate, - candidateReconstruction, - candidateCoefficients, - ref candidateState); + residual, + TransformSize, + this.bitDepth); - if (candidateCost < bestTransformCost) + for (Av1TransformType transformType = Av1TransformType.DctDct; + transformType < Av1TransformType.AllTransformTypes; + transformType++) { - CopyCandidate( + if (!transformType.IsExtendedSetUsed(transformSetType)) + { + continue; + } + + Av1EncoderTransformBlockState candidateState = default; + long candidateCost = this.GetLumaCandidateCost( + writer, + macroBlock, + sourcePlane, + blockOrigin, + prediction, + residual, + bestMode, + selectedAngleDelta, + transformType, + blockContext, + paletteDisabledCost, + largestTransformRate, candidateReconstruction, candidateCoefficients, - reconstructionPlane, - blockOrigin, - retainedCoefficients, - TransformSize, - candidateState, - ref retainedStates[0]); + ref candidateState); - bestTransformCost = candidateCost; + if (candidateCost < bestTransformCost) + { + CopyCandidate( + candidateReconstruction, + candidateCoefficients, + reconstructionPlane, + blockOrigin, + retainedCoefficients, + TransformSize, + candidateState, + ref retainedStates[0]); + + bestTransformCost = candidateCost; + } } } if (this.effort >= 4 && this.picture.Sequence.SequenceHeader.EnableFilterIntra) { - Span filterPrediction = workspace.FilterPrediction; - Span filterResidual = workspace.FilterResidual; - // Each recursive filter prediction and its source residual are independent of transform type. // Prepare them once per filter mode so all legal transforms reuse the same samples. for (Av1FilterIntraMode filterIntraMode = Av1FilterIntraMode.DC; @@ -735,10 +841,10 @@ internal static partial class Av1IntraSuperblockEncoder this.blockWorkspace, sourcePlane, blockOrigin, - filterPrediction, + prediction, above, left, - filterResidual, + residual, filterIntraMode, TransformSize, this.bitDepth); @@ -758,8 +864,8 @@ internal static partial class Av1IntraSuperblockEncoder macroBlock, sourcePlane, blockOrigin, - filterPrediction, - filterResidual, + prediction, + residual, filterIntraMode, transformType, blockContext, @@ -785,12 +891,56 @@ internal static partial class Av1IntraSuperblockEncoder bestMode = Av1PredictionMode.DC; selectedAngleDelta = 0; selectedFilterIntraMode = filterIntraMode; + selectedTransformSize = TransformSize; + } + } + + // Filter-intra mode and transform size form one candidate for RD comparison, just as + // ordinary spatial mode and transform size do in the exhaustive search above. + if (searchEveryTransformSize && + this.picture.Parent.FrameHeader.TransformMode == Av1TransformMode.Select) + { + long splitCost = this.GetSplitLumaCandidateCost( + writer, + macroBlock, + sourcePlane, + reconstructionPlane, + blockOrigin, + tileIndex, + Av1PredictionMode.DC, + 0, + filterIntraMode, + 0, + ReadOnlySpan.Empty, + 0, + paletteDisabledCost, + transformSizeContext, + bestTransformCost, + candidateReconstruction, + candidateCoefficients, + workspace.CandidateTransformBlocks); + + if (splitCost < bestTransformCost) + { + CopySplitCandidate( + candidateReconstruction, + candidateCoefficients, + workspace.CandidateTransformBlocks, + reconstructionPlane, + blockOrigin, + retainedCoefficients, + retainedStates); + + bestTransformCost = splitCost; + bestMode = Av1PredictionMode.DC; + selectedAngleDelta = 0; + selectedFilterIntraMode = filterIntraMode; + selectedTransformSize = Av1TransformSize.Size4x4; } } } } - selectedTransformSize = TransformSize; if (this.effort >= 5 && this.picture.Parent.FrameHeader.AllowScreenContentTools && this.SelectLumaPalette( @@ -817,7 +967,10 @@ internal static partial class Av1IntraSuperblockEncoder selectedFilterIntraMode = Av1FilterIntraMode.AllFilterIntraModes; } + // Efforts six and seven save work by testing transform size only for the global non-palette + // winner. Effort eight and above already tested both sizes inside every candidate. if (this.effort >= 6 && + !searchEveryTransformSize && this.picture.Parent.FrameHeader.TransformMode == Av1TransformMode.Select && paletteInfo.PaletteSizes[0] == 0) { @@ -896,7 +1049,7 @@ internal static partial class Av1IntraSuperblockEncoder TransformSampleCount); Span transformCoefficients = workspace.GetCandidateCoefficients(1)[..TransformSampleCount]; - Span residual = workspace.FilterResidual[..TransformSampleCount]; + Span residual = workspace.Residual[..TransformSampleCount]; Span contexts = workspace.TransformContexts; Span topContexts = contexts[..2]; Span leftContexts = contexts[2..4]; @@ -912,6 +1065,8 @@ internal static partial class Av1IntraSuperblockEncoder TransformSize, useReducedTransformSet); + // Prediction-mode and transform-size symbols belong to the 8x8 coding block, while each + // 4x4 transform contributes its own coefficient rate below. int rate = writer.GetTransformSizeCost(BlockSize, TransformSize, transformSizeContext); if (paletteSize > 0) { @@ -941,6 +1096,9 @@ internal static partial class Av1IntraSuperblockEncoder } long distortion = 0; + + // Raster order is observable here: each retained 4x4 reconstruction supplies reference + // samples and coefficient context to transforms that follow it in the same coding block. for (int transformRow = 0; transformRow < 2; transformRow++) { for (int transformColumn = 0; transformColumn < 2; transformColumn++) @@ -1148,6 +1306,9 @@ internal static partial class Av1IntraSuperblockEncoder int columnOffset = transformColumn * TransformWidth; int modeInfoRow = blockOrigin.Y >> Av1Constants.ModeInfoSizeLog2; int modeInfoColumn = blockOrigin.X >> Av1Constants.ModeInfoSizeLog2; + + // Internal top and left edges come from the candidate mosaic built in raster order. Edges outside + // the 8x8 candidate continue to read committed reconstruction, keeping unsuccessful trials isolated. hasAbove = transformRow > 0 || macroBlock.IsUpAvailable; hasLeft = transformColumn > 0 || macroBlock.IsLeftAvailable; bool rightAvailable = @@ -1291,10 +1452,8 @@ internal static partial class Av1IntraSuperblockEncoder Av1MacroBlockD macroBlock, Buffer2DRegion sourcePlane, Point blockOrigin, - ReadOnlySpan above, - ReadOnlySpan left, - bool hasLeft, - bool hasAbove, + ReadOnlySpan prediction, + ReadOnlySpan residual, Av1PredictionMode mode, int angleDelta, Av1TransformType transformType, @@ -1307,17 +1466,17 @@ internal static partial class Av1IntraSuperblockEncoder { const Av1BlockSize BlockSize = Av1BlockSize.Block8x8; const Av1TransformSize TransformSize = Av1TransformSize.Size8x8; - long distortion = TOperator.EncodeCandidate( + + // Prediction and subtraction were prepared by the owning mode loop. This stage performs only + // transform, quantization, reconstruction, and distortion for the requested transform type. + long distortion = TOperator.EncodePredictionCandidate( this.blockWorkspace, sourcePlane, blockOrigin, + prediction, + residual, candidateReconstruction, - above, - left, - hasLeft, - hasAbove, - mode, - angleDelta, + TransformSize.GetWidth(), candidateCoefficients, TransformSize, transformType, @@ -1328,6 +1487,8 @@ internal static partial class Av1IntraSuperblockEncoder this.bitDepth, ref candidateState); + // Charge every block-level choice that distinguishes this spatial candidate before adding + // coefficient syntax derived from the live neighboring-transform context. int rate = Av1TileWriter.GetLumaModeCost(writer, macroBlock, BlockSize, mode, angleDelta); rate += transformSizeRate; if (mode == Av1PredictionMode.DC) diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.PaletteModeDecision.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.PaletteModeDecision.cs index f959245e7a..31abf04b95 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.PaletteModeDecision.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.PaletteModeDecision.cs @@ -55,6 +55,9 @@ internal static partial class Av1IntraSuperblockEncoder int uniqueColorCount = 0; short minimum = samples[0]; short maximum = samples[0]; + + // An 8x8 block has at most 64 distinct samples, so a compact workspace histogram avoids a + // dictionary allocation while collecting both frequency seeds and range endpoints. foreach (short sample in samples) { int colorIndex = uniqueColors[..uniqueColorCount].IndexOf(sample); @@ -127,7 +130,9 @@ internal static partial class Av1IntraSuperblockEncoder Span centroids = workspace.GetCentroids(0); bool paletteSelected = false; - // Exhaustive ascending size search avoids the reference encoder's speed-dependent pruning. + // Evaluate both frequency-seeded and range-seeded palette families for every legal size. + // Exhaustive ascending size order avoids speed-dependent pruning and gives smaller palettes + // deterministic precedence when complete rate-distortion costs tie. for (int paletteSize = 2; paletteSize <= maximumPaletteSize; paletteSize++) { for (int index = 0; index < paletteSize; index++) @@ -290,6 +295,9 @@ internal static partial class Av1IntraSuperblockEncoder int bitDepth = this.bitDepth.GetBitCount(); int cacheThreshold = 4 << (bitDepth - 8); + + // Nearby colors snap to a coded-neighbor cache entry when the quantization error is bounded. + // Snapping can merge centroids, so sorting and compaction below establish the final coded palette. for (int colorIndex = 0; colorIndex < centroids.Length && !colorCache.IsEmpty; colorIndex++) { int minimumDifference = Math.Abs(centroids[colorIndex] - colorCache[0]); @@ -376,6 +384,8 @@ internal static partial class Av1IntraSuperblockEncoder columns, colorIndexMap); + // Palette prediction and subtraction are invariant for this color map. Reuse them across legal + // transform types, whose enumeration order also supplies deterministic tie precedence. for (Av1TransformType transformType = Av1TransformType.DctDct; transformType < Av1TransformType.AllTransformTypes; transformType++) @@ -420,6 +430,8 @@ internal static partial class Av1IntraSuperblockEncoder long candidateCost = Av1RateDistortion.GetCost(this.rateMultiplier, candidateRate, distortion); if (candidateCost < bestCost) { + // Later palette sizes reuse every candidate span and the shared color map. Publish the + // complete palette state only when this candidate improves the global luma decision. CopyCandidate( candidateReconstruction, candidateCoefficients, diff --git a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs index b15a10b909..41854dd3b5 100644 --- a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs +++ b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs @@ -496,6 +496,9 @@ public class Av1EncoderFrameTests [InlineData(4, true, false, false)] [InlineData(5, true, true, false)] [InlineData(6, true, true, true)] + [InlineData(7, true, true, true)] + [InlineData(8, true, true, true)] + [InlineData(10, true, true, true)] public void EncodeEffortControlsSearchFeatures( int effort, bool enableFilterIntra,