diff --git a/HEIF_IMPLEMENTATION_PLAN.md b/HEIF_IMPLEMENTATION_PLAN.md index 34a62bc84c..e40bce1833 100644 --- a/HEIF_IMPLEMENTATION_PLAN.md +++ b/HEIF_IMPLEMENTATION_PLAN.md @@ -825,13 +825,13 @@ Encoder verification contract: - [~] A production single-tile all-intra writer now walks raster superblocks, analyzes each immediately before entropy coding, reuses one decision workspace and one block workspace, and retains decoder-identical reconstructed references across the tile. Its closed byte and high-bit-depth operators feed the existing superblock boundary without runtime sample-type checks. A byte-exact test compares this composed path with an explicit superblock-then-tile-writer oracle, so producer and writer traversal or coefficient-area drift cannot pass unnoticed. A separate clipped 2x2-superblock regression proves global raster indexing by requiring all four coefficient segments and the bottom-right reconstruction to be populated. Multi-tile ownership and the complete frame/OBU operation remain. - [~] A non-owning encoder-frame view now separates visible conversion regions from coded regions and performs complete left, top, right, bottom, and corner extension across each bordered plane. Current libaom uses 8-sample-aligned coded dimensions, a 32-sample-aligned luma stride with chroma stride derived from it, and a 64-pixel luma border for non-resized all-intra encoding. One operation-ready frame owner now rents the aligned Y, U, and V storage contiguously, exposes non-owning `Buffer2D` plane views, and returns the rent exactly once. A 4K 4:2:0 frame occupies about 13.0 MiB at 8-bit or 26.0 MiB at 10/12-bit; source and reconstruction therefore remain distinct frame owners rather than adding a full-frame copy. The corrected tests use this real ownership path and verify the exact 54 KiB 64x64 4:2:0 rent. The frame-encoder operation now instantiates matching source and reconstruction owners with ordinary `using` lifetimes and converts packed pixels directly into the source owner before extension. - [~] Temporal delimiter, sequence header, frame header, combined-frame tile-group writing, and an internal reduced-still-picture frame operation now exist locally. The remaining required metadata, padding, multi-tile, option, and public encoder paths are not complete. -- [~] Implement superblock and partition analysis for every permitted block size and partition. Efforts zero through eight deliberately split every in-frame node to 8x8 blocks. Effort nine performs recursive live rate-distortion selection at complete 8x8 and 16x16 nodes, while effort ten extends the same search to complete 32x32 nodes. Candidate order matches current libaom: `PARTITION_NONE`, `PARTITION_SPLIT`, `PARTITION_HORZ`, `PARTITION_VERT`, the four asymmetric partitions, then `PARTITION_HORZ_4` and `PARTITION_VERT_4`. Invalid chroma geometries are excluded before evaluation. Each candidate saves and restores the exact partition, coefficient, transform, and palette neighbor edges in one aligned block-workspace owner; trials neither allocate nor copy probability state. Recursive split trials publish each selected child's decoded mode, transform, coefficient, and palette contexts before evaluating its next sibling. Coefficient contexts are published per retained transform rather than broadcasting the first transform over an entire partition leaf. Production picture state retains the compact 8x8 mode allocation below effort nine and explicitly selects 4x4 allocation granularity when sub-8x8 partitions are enabled. Partition selection above 32x32 and effort-dependent pruning remain. +- [~] Implement superblock and partition analysis for every permitted block size and partition. Efforts zero through eight deliberately split every in-frame node to 8x8 blocks. Effort nine performs recursive live rate-distortion selection at complete 8x8 and 16x16 nodes, while effort ten extends the same search to complete 32x32 and 64x64 nodes. Candidate order matches current libaom: `PARTITION_NONE`, `PARTITION_SPLIT`, `PARTITION_HORZ`, `PARTITION_VERT`, the four asymmetric partitions, then `PARTITION_HORZ_4` and `PARTITION_VERT_4`. Invalid chroma geometries are excluded before evaluation. Each candidate saves and restores the exact partition, coefficient, transform, and palette neighbor edges in one aligned block-workspace owner; trials neither allocate nor copy probability state. Recursive split trials publish each selected child's decoded mode, transform, coefficient, and palette contexts before evaluating its next sibling. Coefficient contexts are published per retained transform rather than broadcasting the first transform over an entire partition leaf. Large chroma leaves are evaluated as raster-ordered transform tiles in the existing aligned workspace, and each winning plane is copied to retained storage once. Production picture state retains the compact 8x8 mode allocation below effort nine and explicitly selects 4x4 allocation granularity when sub-8x8 partitions are enabled. Partition selection above 64x64 and effort-dependent pruning remain. - [~] Implement intra mode search, palette, filter intra, chroma-from-luma, and intra-block copy decisions. Live luma search now covers all 13 zero-angle base modes and all six nonzero adjustments for each of the eight directional modes. Joint spatial chroma search covers the same 61 candidates, combines both chroma planes in one rate-distortion decision, and preserves the winning shared angle adjustment. Chroma-from-luma now searches the complete signed alpha alphabet from reconstructed luma and retains its joint U/V syntax. Filter-intra now searches all five predictors after ordinary luma modes. Palette entropy, retained state, production syntax, exhaustive luma and paired chroma palette selection, adaptive screen-content activation, and joint intra-block-copy mode selection are complete. - [ ] Implement inter mode search for bounded sequences, including reference selection and the decoder-supported inter tools. - [~] Current-libaom `av1_quantize_fp_no_qmatrix` arithmetic is implemented as a closed generic forward-quantizer family with Vector512, Vector256, Vector128, and scalar paths, raster-order output, coded 64-point coefficient limits, and scan-order EOB selection. Transform search, coefficient optimization, and lossless behavior remain. -- [~] Implement real rate-distortion selection and make quality and effort change work, size, and output quality. The complete luma and joint chroma candidate sets, including chroma-from-luma, filter-intra, palette, and intra-block copy, now perform live rate-distortion selection. Public quality mapping and effort tiers through exhaustive uniform luma mode/transform search are implemented. Effort nine adds exact recursive 8x8 and 16x16 partition rate-distortion selection, and effort ten extends it through 32x32; partition search above 32x32, effort-dependent pruning, and the remaining sequence searches remain. -- [~] Frame effort now progressively expands the available current search: zero is DC-only, one adds every zero-angle spatial mode, two adds every legal directional adjustment, three refines the preliminary luma winner's transform type, four adds filter-intra and chroma-from-luma, and five adds adaptive palette and intra-block-copy analysis. Lower tiers do not signal unavailable sequence or frame tools, and tiers below five skip the whole-frame screen-content scan. Effort six enables `TX_MODE_SELECT` and compares the winning ordinary spatial or filter-intra luma mode as one 8x8 transform against four raster-ordered 4x4 transforms; each luma palette candidate owns that size comparison from effort six onward. Effort seven searches every legal 8x8 transform type inside every ordinary spatial candidate rather than refining only the preliminary winner. Effort eight also performs the 8x8-versus-four-4x4 comparison inside every ordinary spatial and filter-intra candidate, matching current libaom's per-candidate uniform-transform ownership. Effort nine additionally searches every legal partition at complete 8x8 and 16x16 nodes in current-libaom order, and effort ten extends that recursive search through 32x32. Prediction and residual construction run once per mode and are reused across its legal transform types. Every 4x4 transform searches all legal types with live coefficient contexts and reconstructed intra references. The search reuses the aligned block workspace, preserves only global improvements, and performs no per-block, per-partition, or per-transform rent. Non-skipped intra-block copy writes and costs the current-libaom unsplit variable-transform root; skipped intra-block copy emits no transform-partition symbol. Partition search above 32x32 and effort-dependent model/transform pruning remain. Decoder-visible production cases inspect the emitted restrictions and frame state and decode the produced streams, including real effort-nine streams selecting sub-8x8 and 8x16 rectangular blocks. The complete non-HEVC HEIF/AV1 namespace passes 9,072 of 9,072 through one foreground net11 Release VSTest run. The last independently built `aomdec`, from the then-current `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727` snapshot, accepts the previously generated effort-eight and effort-ten payloads as well as the existing palette and intra-block-copy payloads. The affected encoder, partition, and workspace surface passes 139 of 139 through one foreground net11 Release VSTest run. The net11 Release build and Roslynk compiler and analyzer passes report zero errors. -- [~] Encoder rate accounting converts the entropy writer's live inverse cumulative distributions into current-libaom fixed-point symbol costs without allocating or duplicating probability state. Read-only luma-mode, directional-delta, filter-intra, chroma-mode, block-skip, transform-size, transform-block-skip, and complete transform-coefficient queries share the exact distributions mutated by the subsequent entropy write. Complete coefficient costing follows current libaom's optimized shape: it returns immediately for an empty transform, uses the EOB-specific base-range context, fuses magnitude, sign, base-range, and Golomb accounting into one reverse traversal, and combines repeated full base-range chunks instead of replaying each emitted symbol. Tile-lifetime level and context scratch is reused, the one-coefficient path neither clears nor initializes the forward-neighbor level map, and steady-state queries allocate nothing. Transform-size writing and costing share one subdivision-depth calculation, while shared closed symbol operations keep the writer and cost mappings for transform skip, transform type, and EOB syntax identical without forcing the estimator through the writer's slower two-pass coefficient traversal. The current-libaom fixed-point RD combiner preserves 64-bit distortion and rounds the weighted 1/512-bit rate at the required boundary. Its key-frame multiplier follows libaom's squared DC-quantizer formula and exact 10/12-bit normalization. Live final-block selection evaluates all 61 legal 8x8 luma candidates: the 13 zero-angle base modes in current-libaom order, followed by six nonzero adjustments for each directional mode. Joint chroma selection evaluates the equivalent 61 spatial candidates, combines U and V distortion plus coefficient rate, and charges one live chroma-mode and shared-angle symbol over the actual subsampled 4x4, 4x8, or 8x8 geometry. Chroma-from-luma subsamples the reconstructed luma block once into fixed-stride Q3 stack scratch, subtracts the rounded mean, evaluates all 33 signed alpha values independently for each plane with complete transform RD, and combines the cached plane results across all 1,088 valid joint pairs with one live sign cost and the conditional U/V magnitude costs. This is the allocation-free equivalent of current libaom's exhaustive 33-value path: it requires 66 evaluation transforms rather than transforming every joint pair, preserves DC-before-CfL-before-spatial tie order, and fixes the implicit chroma transform to DCT-DCT. Filter-intra follows ordinary luma candidates, searches all five predictors in syntax order, and evaluates every legal transform while reusing one prepared prediction and source residual per filter mode. Every candidate includes its live mode, angle, filter mode, alpha, and coefficient rate plus normalized pixel-domain distortion. Each prepared reference edge retains the common-corner prefix and twice the transform dimension required by directional prediction. A shared encoder/decoder availability calculation selects reconstructed top-right and bottom-left extensions according to tile, frame, superblock, and block reconstruction order; unavailable extensions repeat the nearest coded endpoint. Missing top or left edges retain current libaom's perpendicular-sample and bit-depth-midpoint rules. Directional prediction applies the AV1 three-degree adjustment step and reuses transform workspace for zone-three transposition before the transform overwrites it, keeping candidate evaluation allocation-free. The winning luma and chroma signed adjustments are retained in the packed final-block state consumed by the tile writer. The tile writer invokes these reusable workspace-backed selectors after mapping current neighbors and immediately before writing each block, so later decisions see reconstructed samples, coefficient contexts, and CDF updates from every preceding block. Block skip is read only after the callback has combined every coded plane. Luma and chroma candidate scratch is partitioned from the encoder's single aligned reusable block workspace; transform-size search uses that owner for four retained 4x4 transform states, local coefficient contexts, and the compact trial reconstruction needed to preserve the best result. No candidate path rents a buffer per block or per transform. Only a newly winning candidate is copied into retained frame storage. Production fixtures force every luma base predictor, both extreme adjustments in all three directional zones, available top-right and bottom-left extensions, high-bit-depth adjustment propagation, exact signed luma and chroma angle-rate terms, joint U/V decisions, packed chroma state, and 4:2:0, 4:2:2, and 4:4:4 transform geometry. The CfL fixtures derive target chroma from a pilot production encode's actual reconstructed luma through an independent scalar Q3 oracle and prove exact positive/negative alpha syntax plus zero-residual DCT-DCT reconstruction for all three subsampling geometries at 8, 10, and 12 bits. The stable fixed-DC traversal comparison uses neutral samples for which both the baseline and live search are contractually DC and skipped, instead of relying on textured content to happen to select the baseline mode. Luma palette selection now evaluates dominant-color and one-dimensional K-means candidates for every legal size, snaps near-cache colors with the reference threshold and tie order, removes duplicate snapped colors, extends boundary maps from active samples, and performs complete transform rate-distortion search. Ordinary DC and filter-intra candidates pay the palette-disabled symbol whenever screen-content syntax is enabled. The exact net11 Release rebuild reports 1,992 test-project warnings and zero errors, all 58 intra-superblock cases pass, all 8,935 AVIF cases pass, and all 230 HEIF cases pass. Remaining mode decision work includes transform-size coverage for filter-intra and palette, broader joint mode/transform refinement, partition search above 32x32 and effort-dependent pruning. Non-empty intra blocks deliberately remain non-skipped, matching current libaom; later inter mode selection owns its distinct skip-transform RD decision. +- [~] Implement real rate-distortion selection and make quality and effort change work, size, and output quality. The complete luma and joint chroma candidate sets, including chroma-from-luma, filter-intra, palette, and intra-block copy, now perform live rate-distortion selection. Public quality mapping and effort tiers through exhaustive uniform luma mode/transform search are implemented. Effort nine adds exact recursive 8x8 and 16x16 partition rate-distortion selection, and effort ten extends it through 64x64; partition search above 64x64, effort-dependent pruning, and the remaining sequence searches remain. +- [~] Frame effort now progressively expands the available current search: zero is DC-only, one adds every zero-angle spatial mode, two adds every legal directional adjustment, three refines the preliminary luma winner's transform type, four adds filter-intra and chroma-from-luma, and five adds adaptive palette and intra-block-copy analysis. Lower tiers do not signal unavailable sequence or frame tools, and tiers below five skip the whole-frame screen-content scan. Effort six enables `TX_MODE_SELECT` and compares the winning ordinary spatial or filter-intra luma mode as one 8x8 transform against four raster-ordered 4x4 transforms; each luma palette candidate owns that size comparison from effort six onward. Effort seven searches every legal 8x8 transform type inside every ordinary spatial candidate rather than refining only the preliminary winner. Effort eight also performs the 8x8-versus-four-4x4 comparison inside every ordinary spatial and filter-intra candidate, matching current libaom's per-candidate uniform-transform ownership. Effort nine additionally searches every legal partition at complete 8x8 and 16x16 nodes in current-libaom order, and effort ten extends that recursive search through 64x64. Prediction and residual construction run once per mode and are reused across its legal transform types. A 64x64 4:4:4 leaf evaluates four 32x32 transforms per chroma plane, retains sparse transform state at coefficient-area offsets, and emits every U transform before every V transform as required by AV1 residual traversal. Every 4x4 transform searches all legal types with live coefficient contexts and reconstructed intra references. The search reuses the aligned block workspace, preserves only global improvements, and performs no per-block, per-partition, or per-transform rent. Non-skipped intra-block copy writes and costs the current-libaom unsplit variable-transform root; skipped intra-block copy emits no transform-partition symbol. Partition search above 64x64 and effort-dependent model/transform pruning remain. Decoder-visible production cases inspect the emitted restrictions and frame state and decode the produced streams, including real effort-nine streams selecting sub-8x8 and 8x16 rectangular blocks. The complete non-HEVC HEIF/AV1 namespace passes 9,074 of 9,074 through one foreground net11 Release VSTest run. The last independently built `aomdec`, from the then-current `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727` snapshot, accepts the previously generated effort-eight and effort-ten payloads as well as the existing palette and intra-block-copy payloads. The affected encoder, partition, and workspace surface passes 139 of 139 through one foreground net11 Release VSTest run. The net11 Release build and Roslynk compiler and analyzer passes report zero errors. +- [~] Encoder rate accounting converts the entropy writer's live inverse cumulative distributions into current-libaom fixed-point symbol costs without allocating or duplicating probability state. Read-only luma-mode, directional-delta, filter-intra, chroma-mode, block-skip, transform-size, transform-block-skip, and complete transform-coefficient queries share the exact distributions mutated by the subsequent entropy write. Complete coefficient costing follows current libaom's optimized shape: it returns immediately for an empty transform, uses the EOB-specific base-range context, fuses magnitude, sign, base-range, and Golomb accounting into one reverse traversal, and combines repeated full base-range chunks instead of replaying each emitted symbol. Tile-lifetime level and context scratch is reused, the one-coefficient path neither clears nor initializes the forward-neighbor level map, and steady-state queries allocate nothing. Transform-size writing and costing share one subdivision-depth calculation, while shared closed symbol operations keep the writer and cost mappings for transform skip, transform type, and EOB syntax identical without forcing the estimator through the writer's slower two-pass coefficient traversal. The current-libaom fixed-point RD combiner preserves 64-bit distortion and rounds the weighted 1/512-bit rate at the required boundary. Its key-frame multiplier follows libaom's squared DC-quantizer formula and exact 10/12-bit normalization. Live final-block selection evaluates all 61 legal 8x8 luma candidates: the 13 zero-angle base modes in current-libaom order, followed by six nonzero adjustments for each directional mode. Joint chroma selection evaluates the equivalent 61 spatial candidates, combines U and V distortion plus coefficient rate, and charges one live chroma-mode and shared-angle symbol over the actual subsampled 4x4, 4x8, or 8x8 geometry. Chroma-from-luma subsamples the reconstructed luma block once into fixed-stride Q3 stack scratch, subtracts the rounded mean, evaluates all 33 signed alpha values independently for each plane with complete transform RD, and combines the cached plane results across all 1,088 valid joint pairs with one live sign cost and the conditional U/V magnitude costs. This is the allocation-free equivalent of current libaom's exhaustive 33-value path: it requires 66 evaluation transforms rather than transforming every joint pair, preserves DC-before-CfL-before-spatial tie order, and fixes the implicit chroma transform to DCT-DCT. Filter-intra follows ordinary luma candidates, searches all five predictors in syntax order, and evaluates every legal transform while reusing one prepared prediction and source residual per filter mode. Every candidate includes its live mode, angle, filter mode, alpha, and coefficient rate plus normalized pixel-domain distortion. Each prepared reference edge retains the common-corner prefix and twice the transform dimension required by directional prediction. A shared encoder/decoder availability calculation selects reconstructed top-right and bottom-left extensions according to tile, frame, superblock, and block reconstruction order; unavailable extensions repeat the nearest coded endpoint. Missing top or left edges retain current libaom's perpendicular-sample and bit-depth-midpoint rules. Directional prediction applies the AV1 three-degree adjustment step and reuses transform workspace for zone-three transposition before the transform overwrites it, keeping candidate evaluation allocation-free. The winning luma and chroma signed adjustments are retained in the packed final-block state consumed by the tile writer. The tile writer invokes these reusable workspace-backed selectors after mapping current neighbors and immediately before writing each block, so later decisions see reconstructed samples, coefficient contexts, and CDF updates from every preceding block. Block skip is read only after the callback has combined every coded plane. Luma and chroma candidate scratch is partitioned from the encoder's single aligned reusable block workspace; transform-size search uses that owner for four retained 4x4 transform states, local coefficient contexts, and the compact trial reconstruction needed to preserve the best result. No candidate path rents a buffer per block or per transform. Only a newly winning candidate is copied into retained frame storage. Production fixtures force every luma base predictor, both extreme adjustments in all three directional zones, available top-right and bottom-left extensions, high-bit-depth adjustment propagation, exact signed luma and chroma angle-rate terms, joint U/V decisions, packed chroma state, and 4:2:0, 4:2:2, and 4:4:4 transform geometry. The CfL fixtures derive target chroma from a pilot production encode's actual reconstructed luma through an independent scalar Q3 oracle and prove exact positive/negative alpha syntax plus zero-residual DCT-DCT reconstruction for all three subsampling geometries at 8, 10, and 12 bits. The stable fixed-DC traversal comparison uses neutral samples for which both the baseline and live search are contractually DC and skipped, instead of relying on textured content to happen to select the baseline mode. Luma palette selection now evaluates dominant-color and one-dimensional K-means candidates for every legal size, snaps near-cache colors with the reference threshold and tie order, removes duplicate snapped colors, extends boundary maps from active samples, and performs complete transform rate-distortion search. Ordinary DC and filter-intra candidates pay the palette-disabled symbol whenever screen-content syntax is enabled. The exact net11 Release rebuild reports 1,992 test-project warnings and zero errors, all 58 intra-superblock cases pass, all 8,935 AVIF cases pass, and all 230 HEIF cases pass. Remaining mode decision work includes transform-size coverage for filter-intra and palette, broader joint mode/transform refinement, partition search above 64x64 and effort-dependent pruning. Non-empty intra blocks deliberately remain non-skipped, matching current libaom; later inter mode selection owns its distinct skip-transform RD decision. - [~] The tile writer now publishes one packed coefficient context per covered 4x4 edge unit and derives luma/chroma skip plus DC-sign contexts from the complete transform edges using current-libaom units. Partition, transform, and coefficient neighbor state retains only the above and left context regions used by current libaom; the unused third top-left region, its granularity state, and its unused sentinel are removed. One picture owner now packs segmentation plus every tile's partition, luma, chroma, and transform edges into one clean byte allocation with typed non-owning views; together with the separately typed packed mode-information owner, the complete picture state uses two allocator rents rather than seven. Exact aligned lengths, clean initialization, and balanced exactly-once returns are covered in Release. Multi-tile payload ownership and verified CDF update behavior remain. - [~] Encoder mode information now uses a frame-owned integer alias grid over a packed 8-byte value allocation, matching current libaom's `mi_grid_base` and `mi_alloc` relationship without a managed object or reference per 4x4 entry. The visible dimensions are aligned to eight luma samples, the grid stride and allocated row count are aligned to 32 mode-information units, and optional 8x8 allocation granularity reduces the value store in both dimensions exactly as current libaom does. One clean ImageSharp byte owner contains both independently typed regions, reducing libaom's two allocation lifetimes to one without a copy. At 4K, the 4x4 layout occupies about 6.0 MiB in total; the 8x8 layout occupies about 3.0 MiB. Exact geometry, clean allocation, typed lengths, aligned mapping, untouched row padding, and exactly-once return pass 4 of 4 direct net11 VSTest cases in Release. Every coded 4x4 cell covered by square, rectangular, or clipped edge blocks maps to its owning allocation entry before context-dependent symbols are written. Packed syntax, relative neighbor lookup, full block mapping, writer traversal, entropy, and OBU coverage pass 1,947 of 1,947 direct net11 VSTest cases in Release; complete mode decision still remains. - [~] The final-block decision workspace uses one reusable 8.3 KiB ImageSharp allocator owner. It contains 1,024 explicitly packed 8-byte final-block entries and the 341 preorder partition bytes required by a complete 128x128-through-8x8 quadtree, replacing separate managed arrays. Palette colors now have their own current-block value and are copied only to the picture edges that later blocks can reference, so enabling palette mode does not add 50 bytes to every final-block entry. Construction and the explicit per-superblock reset initialize every syntax field, including the nonzero sentinel that disables filter-intra prediction; pooled quantizer, prediction, partition, and current-palette bytes cannot leak into the next decision pass. Complete mode decision still remains. diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs index a760ab8852..c2664d4052 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs @@ -18,7 +18,7 @@ internal readonly ref struct Av1EncoderModeDecisionWorkspace /// /// The largest coding-block dimension evaluated directly by the current partition search. /// - public const int MaximumBlockDimension = 32; + public const int MaximumBlockDimension = 64; /// /// The maximum number of samples in one directly evaluated coding block. @@ -30,6 +30,11 @@ internal readonly ref struct Av1EncoderModeDecisionWorkspace /// public const int CandidateTransformBlockCount = 4; + /// + /// The maximum number of transform states needed while evaluating both chroma planes of one 64x64 block. + /// + public const int MaximumCandidateTransformBlockCount = 8; + /// /// The required workspace length in signed-integer storage elements. /// @@ -43,10 +48,10 @@ internal readonly ref struct Av1EncoderModeDecisionWorkspace private const int CandidateCoefficientStorageOffset = CandidateSampleStorageOffset + CandidateSampleStorageLength; private const int CandidateCoefficientStorageLength = 2 * MaximumSampleCount; private const int CandidateTransformBlockStorageOffset = CandidateCoefficientStorageOffset + CandidateCoefficientStorageLength; - private const int CandidateTransformBlockStorageLength = CandidateTransformBlockCount; + private const int CandidateTransformBlockStorageLength = MaximumCandidateTransformBlockCount; private const int TransformContextStorageOffset = CandidateTransformBlockStorageOffset + CandidateTransformBlockStorageLength; private const int TransformContextStorageLength = - 2 * (MaximumBlockDimension >> Av1Constants.ModeInfoSizeLog2) * sizeof(byte) / sizeof(int); + 4 * (MaximumBlockDimension >> Av1Constants.ModeInfoSizeLog2) * sizeof(byte) / sizeof(int); private const int TransientStorageOffset = TransformContextStorageOffset + TransformContextStorageLength; private const int ChromaFromLumaSampleCount = Av1ChromaFromLumaContext.BufferLength; @@ -100,14 +105,14 @@ internal readonly ref struct Av1EncoderModeDecisionWorkspace => new(this.storage[TransientStorageOffset..]); /// - /// Gets the transform state retained while evaluating a uniform 4x4 luma layout. + /// Gets the transform states retained while evaluating a multi-transform candidate. /// public Span CandidateTransformBlocks => MemoryMarshal.Cast( this.storage.Slice(CandidateTransformBlockStorageOffset, CandidateTransformBlockStorageLength)); /// - /// Gets the two above and two left coefficient contexts used by a uniform 4x4 luma layout. + /// Gets coefficient-context edges shared by mutually exclusive luma and chroma transform trials. /// public Span TransformContexts => MemoryMarshal.AsBytes(this.storage.Slice(TransformContextStorageOffset, TransformContextStorageLength)); diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs index d0e648eec2..e454da6e36 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs @@ -52,8 +52,8 @@ internal static partial class Av1IntraSuperblockEncoder Av1TransformSize transformSize, Span retainedBlueCoefficients, Span retainedRedCoefficients, - ref Av1EncoderTransformBlockState retainedBlueState, - ref Av1EncoderTransformBlockState retainedRedState, + Span retainedBlueStates, + Span retainedRedStates, ref Av1EncoderPaletteInfo paletteInfo, out int selectedAngleDelta, out byte selectedChromaFromLumaIndex, @@ -69,6 +69,34 @@ internal static partial class Av1IntraSuperblockEncoder int width = transformSize.GetWidth(); int height = transformSize.GetHeight(); int sampleCount = transformSize.GetSize2d(); + Av1BlockSize chromaBlockSize = blockSize.GetSubsampled( + colorConfig.SubSamplingX, + colorConfig.SubSamplingY); + + if (chromaBlockSize.GetWidth() > width || chromaBlockSize.GetHeight() > height) + { + return this.SelectTiledChromaMode( + writer, + macroBlock, + modeInfo, + lumaOrigin, + chromaOrigin, + blockSize, + chromaBlockSize, + tileIndex, + lumaMode, + transformSize, + retainedBlueCoefficients, + retainedRedCoefficients, + retainedBlueStates, + retainedRedStates, + paletteInfo, + out selectedAngleDelta, + out selectedChromaFromLumaIndex, + out selectedChromaFromLumaSigns, + out selectedCost); + } + int modeInfoRow = lumaOrigin.Y >> Av1Constants.ModeInfoSizeLog2; int modeInfoColumn = lumaOrigin.X >> Av1Constants.ModeInfoSizeLog2; bool hasLeft = macroBlock.IsLeftAvailable; @@ -152,7 +180,6 @@ internal static partial class Av1IntraSuperblockEncoder ReadOnlySpan blueLeft = blueLeftStorage.Slice(1, height * 2); ReadOnlySpan redAbove = redAboveStorage.Slice(1, width * 2); ReadOnlySpan redLeft = redLeftStorage.Slice(1, height * 2); - Av1BlockSize chromaBlockSize = blockSize.GetSubsampled(colorConfig.SubSamplingX, colorConfig.SubSamplingY); Av1TransformBlockContext blueContext = Av1TileWriter.GetTransformBlockContexts( Av1ComponentType.Chroma, this.picture.CbDcSignLevelCoefficientNeighbors[tileIndex], @@ -255,7 +282,7 @@ internal static partial class Av1IntraSuperblockEncoder retainedBlueCoefficients, transformSize, candidateBlueState, - ref retainedBlueState); + ref retainedBlueStates[0]); CopyCandidate( candidateRedReconstruction, @@ -265,7 +292,7 @@ internal static partial class Av1IntraSuperblockEncoder retainedRedCoefficients, transformSize, candidateRedState, - ref retainedRedState); + ref retainedRedStates[0]); bestCost = candidateCost; bestMode = chromaMode; @@ -447,7 +474,7 @@ internal static partial class Av1IntraSuperblockEncoder retainedBlueCoefficients, transformSize, candidateBlueState, - ref retainedBlueState); + ref retainedBlueStates[0]); Av1EncoderTransformBlockState candidateRedState = default; _ = this.GetChromaFromLumaPlaneCost( @@ -474,7 +501,7 @@ internal static partial class Av1IntraSuperblockEncoder retainedRedCoefficients, transformSize, candidateRedState, - ref retainedRedState); + ref retainedRedStates[0]); } } @@ -498,8 +525,8 @@ internal static partial class Av1IntraSuperblockEncoder candidateRedCoefficients[..sampleCount], retainedBlueCoefficients, retainedRedCoefficients, - ref retainedBlueState, - ref retainedRedState, + ref retainedBlueStates[0], + ref retainedRedStates[0], ref bestCost, ref paletteInfo)) { @@ -513,6 +540,408 @@ internal static partial class Av1IntraSuperblockEncoder return bestMode; } + private Av1ChromaPredictionMode SelectTiledChromaMode( + Av1SymbolEncoder writer, + Av1MacroBlockD macroBlock, + Av1MacroBlockModeInfo modeInfo, + Point lumaOrigin, + Point chromaOrigin, + Av1BlockSize blockSize, + Av1BlockSize chromaBlockSize, + ushort tileIndex, + Av1PredictionMode lumaMode, + Av1TransformSize transformSize, + Span retainedBlueCoefficients, + Span retainedRedCoefficients, + Span retainedBlueStates, + Span retainedRedStates, + Av1EncoderPaletteInfo paletteInfo, + out int selectedAngleDelta, + out byte selectedChromaFromLumaIndex, + out sbyte selectedChromaFromLumaSigns, + out long selectedCost) + { + Av1EncoderModeDecisionWorkspace workspace = + this.blockWorkspace.GetModeDecisionWorkspace(); + + ObuColorConfig colorConfig = this.picture.Sequence.SequenceHeader.ColorConfig; + int subsamplingX = colorConfig.SubSamplingX ? 1 : 0; + int subsamplingY = colorConfig.SubSamplingY ? 1 : 0; + int blockWidth = chromaBlockSize.GetWidth(); + int blockHeight = chromaBlockSize.GetHeight(); + int blockSampleCount = blockWidth * blockHeight; + int transformWidth = transformSize.GetWidth(); + int transformHeight = transformSize.GetHeight(); + int transformColumnCount = blockWidth / transformWidth; + int transformRowCount = blockHeight / transformHeight; + int transformBlockCount = transformColumnCount * transformRowCount; + Span candidateBlueReconstruction = + workspace.GetCandidateReconstruction(0)[..blockSampleCount]; + + Span candidateRedReconstruction = + workspace.GetCandidateReconstruction(1)[..blockSampleCount]; + + Span candidateBlueCoefficients = + workspace.GetCandidateCoefficients(0)[..blockSampleCount]; + + Span candidateRedCoefficients = + workspace.GetCandidateCoefficients(1)[..blockSampleCount]; + + Span candidateStates = workspace.CandidateTransformBlocks; + Span candidateBlueStates = + candidateStates[..transformBlockCount]; + + Span candidateRedStates = + candidateStates.Slice(transformBlockCount, transformBlockCount); + + int contextWidth = chromaBlockSize.Get4x4WideCount(); + int contextHeight = chromaBlockSize.Get4x4HighCount(); + Span contexts = workspace.TransformContexts; + Span blueTopContexts = contexts[..contextWidth]; + Span blueLeftContexts = contexts.Slice(contextWidth, contextHeight); + Span redTopContexts = contexts.Slice(contextWidth + contextHeight, contextWidth); + Span redLeftContexts = contexts.Slice( + (2 * contextWidth) + contextHeight, + contextHeight); + + Av1NeighborArrayUnit blueNeighbors = + this.picture.CbDcSignLevelCoefficientNeighbors[tileIndex]; + + Av1NeighborArrayUnit redNeighbors = + this.picture.CrDcSignLevelCoefficientNeighbors[tileIndex]; + + int blueTopIndex = blueNeighbors.GetTopIndex(chromaOrigin); + int blueLeftIndex = blueNeighbors.GetLeftIndex(chromaOrigin); + int redTopIndex = redNeighbors.GetTopIndex(chromaOrigin); + int redLeftIndex = redNeighbors.GetLeftIndex(chromaOrigin); + Buffer2DRegion blueSource = this.source.GetPlane(Av1Plane.U); + Buffer2DRegion redSource = this.source.GetPlane(Av1Plane.V); + Buffer2DRegion blueReconstruction = this.reconstruction.GetPlane(Av1Plane.U); + Buffer2DRegion redReconstruction = this.reconstruction.GetPlane(Av1Plane.V); + bool hasLumaPalette = paletteInfo.PaletteSizes[0] != 0; + int paletteDisabledCost = this.picture.Parent.FrameHeader.AllowScreenContentTools + ? writer.GetPaletteUvModeCost(false, hasLumaPalette) + : 0; + + int baseModeCount = ChromaModeSearchOrder.Length; + int deltaCount = AngleDeltaSearchOrder.Length; + int directionalModeCount = + (int)Av1ChromaPredictionMode.Directional67Degrees - + (int)Av1ChromaPredictionMode.Vertical + + 1; + + int candidateCount = baseModeCount + (directionalModeCount * deltaCount); + long bestCost = long.MaxValue; + Av1ChromaPredictionMode bestMode = Av1ChromaPredictionMode.DC; + selectedAngleDelta = 0; + selectedChromaFromLumaIndex = 0; + selectedChromaFromLumaSigns = 0; + + // A large chroma block is predicted and transformed in the same raster order used by the tile + // writer. Each completed transform supplies both reconstructed edges and coefficient contexts. + for (int candidateIndex = 0; candidateIndex < candidateCount; candidateIndex++) + { + Av1ChromaPredictionMode chromaMode; + int angleDelta; + if (candidateIndex < baseModeCount) + { + chromaMode = ChromaModeSearchOrder[candidateIndex]; + angleDelta = 0; + } + else + { + int adjustedIndex = candidateIndex - baseModeCount; + chromaMode = (Av1ChromaPredictionMode)( + (int)Av1ChromaPredictionMode.Vertical + (adjustedIndex / deltaCount)); + + angleDelta = AngleDeltaSearchOrder[adjustedIndex % deltaCount]; + } + + blueNeighbors.Top.Slice(blueTopIndex, contextWidth).CopyTo(blueTopContexts); + blueNeighbors.Left.Slice(blueLeftIndex, contextHeight).CopyTo(blueLeftContexts); + redNeighbors.Top.Slice(redTopIndex, contextWidth).CopyTo(redTopContexts); + redNeighbors.Left.Slice(redLeftIndex, contextHeight).CopyTo(redLeftContexts); + Av1PredictionMode predictionMode = chromaMode.ToLumaMode(); + long distortion = this.GetTiledChromaPlaneCost( + writer, + macroBlock, + lumaOrigin, + chromaOrigin, + blockSize, + chromaBlockSize, + transformSize, + subsamplingX, + subsamplingY, + lumaMode, + predictionMode, + angleDelta, + Av1Plane.U, + blueSource, + blueReconstruction, + candidateBlueReconstruction, + candidateBlueCoefficients, + candidateBlueStates, + blueTopContexts, + blueLeftContexts, + out int blueRate); + + distortion += this.GetTiledChromaPlaneCost( + writer, + macroBlock, + lumaOrigin, + chromaOrigin, + blockSize, + chromaBlockSize, + transformSize, + subsamplingX, + subsamplingY, + lumaMode, + predictionMode, + angleDelta, + Av1Plane.V, + redSource, + redReconstruction, + candidateRedReconstruction, + candidateRedCoefficients, + candidateRedStates, + redTopContexts, + redLeftContexts, + out int redRate); + + int rate = Av1TileWriter.GetChromaModeCost( + writer, + this.picture.Parent.FrameHeader, + colorConfig, + modeInfo, + blockSize, + lumaMode, + chromaMode, + angleDelta); + + rate += blueRate + redRate; + if (chromaMode == Av1ChromaPredictionMode.DC) + { + rate += paletteDisabledCost; + } + + long candidateCost = Av1RateDistortion.GetCost(this.rateMultiplier, rate, distortion); + if (candidateCost < bestCost) + { + CopyTiledCandidate( + candidateBlueReconstruction, + candidateBlueCoefficients, + candidateBlueStates, + blueReconstruction, + chromaOrigin, + blockWidth, + blockHeight, + transformSize, + retainedBlueCoefficients, + retainedBlueStates); + + CopyTiledCandidate( + candidateRedReconstruction, + candidateRedCoefficients, + candidateRedStates, + redReconstruction, + chromaOrigin, + blockWidth, + blockHeight, + transformSize, + retainedRedCoefficients, + retainedRedStates); + + bestCost = candidateCost; + bestMode = chromaMode; + selectedAngleDelta = angleDelta; + } + } + + selectedCost = bestCost; + return bestMode; + } + + private long GetTiledChromaPlaneCost( + Av1SymbolEncoder writer, + Av1MacroBlockD macroBlock, + Point lumaOrigin, + Point chromaOrigin, + Av1BlockSize blockSize, + Av1BlockSize chromaBlockSize, + Av1TransformSize transformSize, + int subsamplingX, + int subsamplingY, + Av1PredictionMode lumaMode, + Av1PredictionMode predictionMode, + int angleDelta, + Av1Plane plane, + Buffer2DRegion source, + Buffer2DRegion reconstruction, + Span candidateReconstruction, + Span candidateCoefficients, + Span candidateStates, + Span topContexts, + Span leftContexts, + out int rate) + { + Av1EncoderModeDecisionWorkspace workspace = + this.blockWorkspace.GetModeDecisionWorkspace(); + + int blockWidth = chromaBlockSize.GetWidth(); + int blockHeight = chromaBlockSize.GetHeight(); + int transformWidth = transformSize.GetWidth(); + int transformHeight = transformSize.GetHeight(); + int transformSampleCount = transformSize.GetSize2d(); + int transformWidth4x4 = transformSize.Get4x4WideCount(); + int transformHeight4x4 = transformSize.Get4x4HighCount(); + Av1TransformType transformType = Av1SymbolContextHelper.GetDefaultIntraTransformType( + predictionMode, + transformSize, + this.picture.Parent.FrameHeader.UseReducedTransformSet); + + Av1ComponentType componentType = Av1ComponentType.Chroma; + Span prediction = workspace.Prediction[..transformSampleCount]; + Span residual = workspace.Residual[..transformSampleCount]; + Span aboveStorage = workspace.GetReferenceSamples(0); + Span leftStorage = workspace.GetReferenceSamples(1); + int coefficientOffset = 0; + int transformIndex = 0; + long distortion = 0; + rate = 0; + for (int transformRow = 0; transformRow < blockHeight / transformHeight; transformRow++) + { + int rowOffset = transformRow * transformHeight; + for (int transformColumn = 0; transformColumn < blockWidth / transformWidth; transformColumn++) + { + int columnOffset = transformColumn * transformWidth; + int reconstructionOffset = (rowOffset * blockWidth) + columnOffset; + Point transformOrigin = chromaOrigin + new Size(columnOffset, rowOffset); + this.PrepareTransformReferenceSamples( + reconstruction, + lumaOrigin, + chromaOrigin, + blockSize, + macroBlock, + transformRow, + transformColumn, + blockWidth, + transformSize, + subsamplingX, + subsamplingY, + candidateReconstruction, + aboveStorage, + leftStorage, + out bool hasLeft, + out bool hasAbove); + + TOperator.PrepareIntra( + this.blockWorkspace, + source, + transformOrigin, + prediction, + aboveStorage.Slice(1, transformWidth * 2), + leftStorage.Slice(1, transformHeight * 2), + hasLeft, + hasAbove, + predictionMode, + angleDelta, + residual, + transformSize, + this.bitDepth); + + Av1TransformBlockContext blockContext = Av1TileWriter.GetTransformBlockContexts( + componentType, + topContexts.Slice(transformColumn * transformWidth4x4, transformWidth4x4), + leftContexts.Slice(transformRow * transformHeight4x4, transformHeight4x4), + chromaBlockSize, + transformSize); + + Span transformCoefficients = candidateCoefficients.Slice( + coefficientOffset, + transformSampleCount); + + ref Av1EncoderTransformBlockState state = ref candidateStates[transformIndex++]; + distortion += TOperator.EncodePredictionCandidate( + this.blockWorkspace, + source, + transformOrigin, + prediction, + residual, + candidateReconstruction[reconstructionOffset..], + blockWidth, + transformCoefficients, + transformSize, + transformType, + plane, + this.quantization.QIndex[0], + this.quantization.DeltaQDc[(int)plane], + this.quantization.DeltaQAc[(int)plane], + this.bitDepth, + ref state); + + rate += writer.GetCoefficientCost( + transformSize, + transformType, + lumaMode, + transformCoefficients, + componentType, + blockContext, + state.EndOfBlock, + this.picture.Parent.FrameHeader.UseReducedTransformSet, + Av1FilterIntraMode.AllFilterIntraModes, + usesInterTransformSet: false); + + byte coefficientContext = Av1SymbolContextHelper.GetCoefficientContext( + transformCoefficients, + transformSize, + transformType, + state.EndOfBlock); + + topContexts + .Slice(transformColumn * transformWidth4x4, transformWidth4x4) + .Fill(coefficientContext); + + leftContexts + .Slice(transformRow * transformHeight4x4, transformHeight4x4) + .Fill(coefficientContext); + + coefficientOffset += transformSampleCount; + } + } + + return distortion; + } + + private static void CopyTiledCandidate( + ReadOnlySpan candidateReconstruction, + ReadOnlySpan candidateCoefficients, + ReadOnlySpan candidateStates, + Buffer2DRegion reconstruction, + Point blockOrigin, + int blockWidth, + int blockHeight, + Av1TransformSize transformSize, + Span retainedCoefficients, + Span retainedStates) + { + int blockSampleCount = blockWidth * blockHeight; + int transformStateStride = + transformSize.GetSize2d() / + Av1EncoderCoefficientBuffer.TransformBlockUnitCoefficientCount; + + candidateCoefficients[..blockSampleCount].CopyTo(retainedCoefficients); + for (int transformIndex = 0; transformIndex < candidateStates.Length; transformIndex++) + { + retainedStates[transformIndex * transformStateStride] = candidateStates[transformIndex]; + } + + for (int row = 0; row < blockHeight; row++) + { + candidateReconstruction.Slice(row * blockWidth, blockWidth) + .CopyTo(reconstruction.DangerousGetRowSpan(blockOrigin.Y + row).Slice(blockOrigin.X, blockWidth)); + } + } + private long GetChromaFromLumaPlaneCost( Av1SymbolEncoder writer, Av1PredictionMode lumaMode, diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs index d1a15998bc..5f139a164f 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs @@ -184,7 +184,7 @@ internal static partial class Av1IntraSuperblockEncoder Av1PartitionType preparedPartition) { bool searchPartition = blockSize is Av1BlockSize.Block8x8 or Av1BlockSize.Block16x16 || - (this.effort == 10 && blockSize == Av1BlockSize.Block32x32); + (this.effort == 10 && blockSize is Av1BlockSize.Block32x32 or Av1BlockSize.Block64x64); if (!searchPartition) { @@ -652,6 +652,10 @@ internal static partial class Av1IntraSuperblockEncoder colorConfig.SubSamplingX, colorConfig.SubSamplingY); + Av1BlockSize chromaBlockSize = blockSize.GetSubsampled( + colorConfig.SubSamplingX, + colorConfig.SubSamplingY); + Span blueCoefficients = this.coefficientBuffer.GetPlaneSpan(this.superblock.Index, Av1Plane.U); Span redCoefficients = this.coefficientBuffer.GetPlaneSpan(this.superblock.Index, Av1Plane.V); Span blueTransformBlocks = @@ -662,8 +666,8 @@ internal static partial class Av1IntraSuperblockEncoder int chromaTransformIndex = this.codedAreaChroma / Av1EncoderCoefficientBuffer.TransformBlockUnitCoefficientCount; - ref Av1EncoderTransformBlockState blueState = ref blueTransformBlocks[chromaTransformIndex]; - ref Av1EncoderTransformBlockState redState = ref redTransformBlocks[chromaTransformIndex]; + Span retainedBlueStates = blueTransformBlocks[chromaTransformIndex..]; + Span retainedRedStates = redTransformBlocks[chromaTransformIndex..]; long chromaCost = 0; if (block.HasChroma) { @@ -679,8 +683,8 @@ internal static partial class Av1IntraSuperblockEncoder chromaTransformSize, blueCoefficients[this.codedAreaChroma..], redCoefficients[this.codedAreaChroma..], - ref blueState, - ref redState, + retainedBlueStates, + retainedRedStates, ref paletteInfo, out int chromaAngleDelta, out byte chromaFromLumaIndex, @@ -694,16 +698,26 @@ internal static partial class Av1IntraSuperblockEncoder // Skip suppresses coefficient syntax for the entire coding block, not one plane independently. // Preserve normal coefficient coding when any selected luma or chroma transform is nonempty. - bool allTransformsEmpty = lumaTransformEmpty && - (!block.HasChroma || (blueState.EndOfBlock == 0 && redState.EndOfBlock == 0)); + int chromaTransformSampleCount = chromaTransformSize.GetSize2d(); + int chromaTransformBlockCount = + (chromaBlockSize.GetWidth() * chromaBlockSize.GetHeight()) / chromaTransformSampleCount; + + int chromaStateStride = + chromaTransformSampleCount / Av1EncoderCoefficientBuffer.TransformBlockUnitCoefficientCount; + + bool chromaTransformsEmpty = true; + for (int transformIndex = 0; transformIndex < chromaTransformBlockCount; transformIndex++) + { + int stateIndex = transformIndex * chromaStateStride; + chromaTransformsEmpty &= retainedBlueStates[stateIndex].EndOfBlock == 0 && + retainedRedStates[stateIndex].EndOfBlock == 0; + } + + bool allTransformsEmpty = lumaTransformEmpty && (!block.HasChroma || chromaTransformsEmpty); int regularEmptyTransformRate = 0; if (allTransformsEmpty) { - Av1BlockSize chromaBlockSize = blockSize.GetSubsampled( - colorConfig.SubSamplingX, - colorConfig.SubSamplingY); - regularEmptyTransformRate = this.GetEmptyTransformRate( writer, this.picture.LuminanceDcSignLevelCoefficientNeighbors[tileIndex], @@ -769,10 +783,6 @@ internal static partial class Av1IntraSuperblockEncoder this.codedAreaLuma += blockSize.GetWidth() * blockSize.GetHeight(); if (block.HasChroma) { - Av1BlockSize chromaBlockSize = blockSize.GetSubsampled( - colorConfig.SubSamplingX, - colorConfig.SubSamplingY); - this.codedAreaChroma += chromaBlockSize.GetWidth() * chromaBlockSize.GetHeight(); } } @@ -958,8 +968,10 @@ internal static partial class Av1IntraSuperblockEncoder int transformWidth = transformSize.GetWidth(); int transformHeight = transformSize.GetHeight(); int transformSampleCount = transformSize.GetSize2d(); - int transformIndex = 0; + int transformStateOffset = 0; int coefficientOffset = 0; + int transformStateStride = + transformSampleCount / Av1EncoderCoefficientBuffer.TransformBlockUnitCoefficientCount; // Uniform transform blocks are retained and written in raster order. Publishing that same tiling // preserves the distinct top and left contexts consumed by the next coding block in a dry run. @@ -967,7 +979,7 @@ internal static partial class Av1IntraSuperblockEncoder { for (int column = 0; column < blockWidth; column += transformWidth) { - Av1EncoderTransformBlockState state = states[transformIndex++]; + Av1EncoderTransformBlockState state = states[transformStateOffset]; byte context = Av1SymbolContextHelper.GetCoefficientContext( coefficients[coefficientOffset..], transformSize, @@ -981,6 +993,7 @@ internal static partial class Av1IntraSuperblockEncoder EdgeMask); coefficientOffset += transformSampleCount; + transformStateOffset += transformStateStride; } } } @@ -1685,7 +1698,9 @@ internal static partial class Av1IntraSuperblockEncoder } if (this.effort >= 4 && - this.picture.Sequence.SequenceHeader.EnableFilterIntra) + this.picture.Sequence.SequenceHeader.EnableFilterIntra && + blockWidth <= 32 && + blockHeight <= 32) { // Each recursive filter prediction and its source residual are independent of transform type. // Prepare them once per filter mode so all legal transforms reuse the same samples. @@ -2004,12 +2019,18 @@ internal static partial class Av1IntraSuperblockEncoder { Span aboveStorage = workspace.GetReferenceSamples(0); Span leftStorage = workspace.GetReferenceSamples(1); - this.PrepareSplitLumaReferenceSamples( + this.PrepareTransformReferenceSamples( reconstructionPlane, blockOrigin, + blockOrigin, + BlockSize, macroBlock, transformRow, transformColumn, + BlockWidth, + TransformSize, + 0, + 0, candidateReconstruction, aboveStorage, leftStorage, @@ -2167,81 +2188,89 @@ internal static partial class Av1IntraSuperblockEncoder return Av1RateDistortion.GetCost(this.rateMultiplier, rate, distortion); } - private void PrepareSplitLumaReferenceSamples( + private void PrepareTransformReferenceSamples( Buffer2DRegion reconstructionPlane, - Point blockOrigin, + Point lumaBlockOrigin, + Point planeBlockOrigin, + Av1BlockSize blockSize, Av1MacroBlockD macroBlock, int transformRow, int transformColumn, + int planeBlockWidth, + Av1TransformSize transformSize, + int subsamplingX, + int subsamplingY, ReadOnlySpan candidateReconstruction, Span aboveStorage, Span leftStorage, out bool hasLeft, out bool hasAbove) { - const Av1BlockSize BlockSize = Av1BlockSize.Block8x8; - const Av1TransformSize TransformSize = Av1TransformSize.Size4x4; - const int BlockWidth = 8; - const int TransformWidth = 4; - int rowOffset = transformRow * TransformWidth; - int columnOffset = transformColumn * TransformWidth; - int modeInfoRow = blockOrigin.Y >> Av1Constants.ModeInfoSizeLog2; - int modeInfoColumn = blockOrigin.X >> Av1Constants.ModeInfoSizeLog2; + int transformWidth = transformSize.GetWidth(); + int transformHeight = transformSize.GetHeight(); + int rowOffset = transformRow * transformHeight; + int columnOffset = transformColumn * transformWidth; + int modeInfoRow = lumaBlockOrigin.Y >> Av1Constants.ModeInfoSizeLog2; + int modeInfoColumn = lumaBlockOrigin.X >> Av1Constants.ModeInfoSizeLog2; // Internal top and left edges come from the candidate mosaic built in raster order. Edges outside - // the 8x8 candidate continue to read committed reconstruction, keeping unsuccessful trials isolated. + // the candidate continue to read committed reconstruction, keeping unsuccessful trials isolated. hasAbove = transformRow > 0 || macroBlock.IsUpAvailable; hasLeft = transformColumn > 0 || macroBlock.IsLeftAvailable; + int transformRow4x4 = rowOffset >> Av1Constants.ModeInfoSizeLog2; + int transformColumn4x4 = columnOffset >> Av1Constants.ModeInfoSizeLog2; bool rightAvailable = - modeInfoColumn + transformColumn + TransformSize.Get4x4WideCount() < + modeInfoColumn + + ((transformColumn4x4 + transformSize.Get4x4WideCount()) << subsamplingX) < macroBlock.Tile.ModeInfoColumnEnd; bool bottomAvailable = - modeInfoRow + transformRow + TransformSize.Get4x4HighCount() < + modeInfoRow + + ((transformRow4x4 + transformSize.Get4x4HighCount()) << subsamplingY) < macroBlock.Tile.ModeInfoRowEnd; bool hasTopRight = Av1IntraReferenceAvailability.HasTopRight( this.picture.Sequence.SequenceHeader.SuperblockSize, - BlockSize, + blockSize, modeInfoRow, modeInfoColumn, hasAbove, rightAvailable, Av1PartitionType.None, - TransformSize, - transformRow, - transformColumn, - 0, - 0); + transformSize, + transformRow4x4, + transformColumn4x4, + subsamplingX, + subsamplingY); bool hasBottomLeft = Av1IntraReferenceAvailability.HasBottomLeft( this.picture.Sequence.SequenceHeader.SuperblockSize, - BlockSize, + blockSize, modeInfoRow, modeInfoColumn, bottomAvailable, hasLeft, Av1PartitionType.None, - TransformSize, - transformRow, - transformColumn, - 0, - 0); + transformSize, + transformRow4x4, + transformColumn4x4, + subsamplingX, + subsamplingY); - Span above = aboveStorage.Slice(1, TransformWidth * 2); - Span left = leftStorage.Slice(1, TransformWidth * 2); + Span above = aboveStorage.Slice(1, transformWidth * 2); + Span left = leftStorage.Slice(1, transformHeight * 2); if (hasAbove) { if (transformRow > 0) { candidateReconstruction - .Slice(((rowOffset - 1) * BlockWidth) + columnOffset, TransformWidth) + .Slice(((rowOffset - 1) * planeBlockWidth) + columnOffset, transformWidth) .CopyTo(above); } else { - reconstructionPlane.DangerousGetRowSpan(blockOrigin.Y - 1) - .Slice(blockOrigin.X + columnOffset, TransformWidth) + reconstructionPlane.DangerousGetRowSpan(planeBlockOrigin.Y - 1) + .Slice(planeBlockOrigin.X + columnOffset, transformWidth) .CopyTo(above); } } @@ -2250,17 +2279,18 @@ internal static partial class Av1IntraSuperblockEncoder { if (transformColumn > 0) { - for (int row = 0; row < TransformWidth; row++) + for (int row = 0; row < transformHeight; row++) { - left[row] = candidateReconstruction[((rowOffset + row) * BlockWidth) + columnOffset - 1]; + left[row] = candidateReconstruction[ + ((rowOffset + row) * planeBlockWidth) + columnOffset - 1]; } } else { - for (int row = 0; row < TransformWidth; row++) + for (int row = 0; row < transformHeight; row++) { left[row] = reconstructionPlane - .DangerousGetRowSpan(blockOrigin.Y + rowOffset + row)[blockOrigin.X - 1]; + .DangerousGetRowSpan(planeBlockOrigin.Y + rowOffset + row)[planeBlockOrigin.X - 1]; } } } @@ -2268,12 +2298,12 @@ internal static partial class Av1IntraSuperblockEncoder int midpoint = 128 << (this.bitDepth.GetBitCount() - 8); if (!hasAbove) { - above[..TransformWidth].Fill(hasLeft ? left[0] : TOperator.CreateSample(midpoint - 1)); + above[..transformWidth].Fill(hasLeft ? left[0] : TOperator.CreateSample(midpoint - 1)); } if (!hasLeft) { - left[..TransformWidth].Fill(hasAbove ? above[0] : TOperator.CreateSample(midpoint + 1)); + left[..transformHeight].Fill(hasAbove ? above[0] : TOperator.CreateSample(midpoint + 1)); } if (hasTopRight) @@ -2282,42 +2312,42 @@ internal static partial class Av1IntraSuperblockEncoder { candidateReconstruction .Slice( - ((rowOffset - 1) * BlockWidth) + columnOffset + TransformWidth, - TransformWidth) - .CopyTo(above[TransformWidth..]); + ((rowOffset - 1) * planeBlockWidth) + columnOffset + transformWidth, + transformWidth) + .CopyTo(above[transformWidth..]); } else { - reconstructionPlane.DangerousGetRowSpan(blockOrigin.Y - 1) - .Slice(blockOrigin.X + columnOffset + TransformWidth, TransformWidth) - .CopyTo(above[TransformWidth..]); + reconstructionPlane.DangerousGetRowSpan(planeBlockOrigin.Y - 1) + .Slice(planeBlockOrigin.X + columnOffset + transformWidth, transformWidth) + .CopyTo(above[transformWidth..]); } } else { - above[TransformWidth..].Fill(above[TransformWidth - 1]); + above[transformWidth..].Fill(above[transformWidth - 1]); } if (hasBottomLeft) { - for (int row = TransformWidth; row < TransformWidth * 2; row++) + for (int row = transformHeight; row < transformHeight * 2; row++) { left[row] = reconstructionPlane - .DangerousGetRowSpan(blockOrigin.Y + rowOffset + row)[blockOrigin.X - 1]; + .DangerousGetRowSpan(planeBlockOrigin.Y + rowOffset + row)[planeBlockOrigin.X - 1]; } } else { - left[TransformWidth..].Fill(left[TransformWidth - 1]); + left[transformHeight..].Fill(left[transformHeight - 1]); } // Only an interior transform corner belongs to decision scratch. Boundary corners continue // to read the already reconstructed neighboring block so candidate trials remain isolated. TSample corner = hasAbove && hasLeft ? transformRow > 0 && transformColumn > 0 - ? candidateReconstruction[((rowOffset - 1) * BlockWidth) + columnOffset - 1] - : reconstructionPlane.DangerousGetRowSpan(blockOrigin.Y + rowOffset - 1)[ - blockOrigin.X + columnOffset - 1] + ? candidateReconstruction[((rowOffset - 1) * planeBlockWidth) + columnOffset - 1] + : reconstructionPlane.DangerousGetRowSpan(planeBlockOrigin.Y + rowOffset - 1)[ + planeBlockOrigin.X + columnOffset - 1] : hasAbove ? above[0] : hasLeft diff --git a/src/ImageSharp/Formats/Heif/Av1/Tiling/Av1TileWriter.cs b/src/ImageSharp/Formats/Heif/Av1/Tiling/Av1TileWriter.cs index 364e2b7940..45b4b7a3bf 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Tiling/Av1TileWriter.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Tiling/Av1TileWriter.cs @@ -1914,84 +1914,78 @@ internal partial class Av1TileWriter int maximumUnitBlocksWide = Math.Min(maximumUnitBlockSize.Get4x4WideCount(), maximumBlocksWide); int maximumUnitBlocksHigh = Math.Min(maximumUnitBlockSize.Get4x4HighCount(), maximumBlocksHigh); - // Chroma follows the same bounded-region order after scaling both the block and frame edges to its plane. - for (int regionRow = 0; regionRow < maximumBlocksHigh; regionRow += maximumUnitBlocksHigh) + int codedAreaStart = entropyCodingContext.CodedAreaSuperblockUv; + int codedAreaEnd = codedAreaStart; + + // AV1 completes every transform in one chroma plane before advancing to the other plane. + // Both planes use the same coded-area positions because their transform geometry is identical. + for (int planeIndex = 0; planeIndex < 2; planeIndex++) { - int unitHeight = Math.Min(maximumUnitBlocksHigh + regionRow, maximumBlocksHigh); - for (int regionColumn = 0; regionColumn < maximumBlocksWide; regionColumn += maximumUnitBlocksWide) + bool isBluePlane = planeIndex == 0; + Span planeCoefficients = isBluePlane ? blueCoefficients : redCoefficients; + Span planeTransformBlocks = + isBluePlane ? blueTransformBlocks : redTransformBlocks; + + Av1NeighborArrayUnit coefficientNeighbors = + isBluePlane ? cb_dc_sign_level_coeff_na : cr_dc_sign_level_coeff_na; + + int codedArea = codedAreaStart; + for (int regionRow = 0; regionRow < maximumBlocksHigh; regionRow += maximumUnitBlocksHigh) { - int unitWidth = Math.Min(maximumUnitBlocksWide + regionColumn, maximumBlocksWide); - for (int blockRow = regionRow; blockRow < unitHeight; blockRow += transformBlockHeight) + int unitHeight = Math.Min(maximumUnitBlocksHigh + regionRow, maximumBlocksHigh); + for (int regionColumn = 0; regionColumn < maximumBlocksWide; regionColumn += maximumUnitBlocksWide) { - for (int blockColumn = regionColumn; blockColumn < unitWidth; blockColumn += transformBlockWidth) + int unitWidth = Math.Min(maximumUnitBlocksWide + regionColumn, maximumBlocksWide); + for (int blockRow = regionRow; blockRow < unitHeight; blockRow += transformBlockHeight) { - int transformStateIndex = entropyCodingContext.CodedAreaSuperblockUv / - Av1EncoderCoefficientBuffer.TransformBlockUnitCoefficientCount; - ref Av1EncoderTransformBlockState blueTransformBlock = ref blueTransformBlocks[transformStateIndex]; - ref Av1EncoderTransformBlockState redTransformBlock = ref redTransformBlocks[transformStateIndex]; - Point chromaOrigin = chromaBlockOrigin + new Size( - blockColumn << Av1Constants.ModeInfoSizeLog2, - blockRow << Av1Constants.ModeInfoSizeLog2); - - // U and V share transform geometry and type while retaining independent coefficient and EOB state. - Span coefficients = blueCoefficients[entropyCodingContext.CodedAreaSuperblockUv..]; - Av1TransformBlockContext blockContext = GetTransformBlockContexts( - Av1ComponentType.Chroma, - cb_dc_sign_level_coeff_na, - chromaOrigin, - chromaBlockSize, - chromaTransformSize); - - Av1TransformType chromaTransformType = blueTransformBlock.TransformType; - int culLevelCb = writer.WriteCoefficients( - chromaTransformSize, - chromaTransformType, - intraLumaDir, - coefficients, - Av1ComponentType.Chroma, - blockContext, - blueTransformBlock.EndOfBlock, - frameHeader.UseReducedTransformSet, - blk_ptr.FilterIntraMode, - usesInterTransformSet); - - coefficients = redCoefficients[entropyCodingContext.CodedAreaSuperblockUv..]; - blockContext = GetTransformBlockContexts( - Av1ComponentType.Chroma, - cr_dc_sign_level_coeff_na, - chromaOrigin, - chromaBlockSize, - chromaTransformSize); - - int culLevelCr = writer.WriteCoefficients( - chromaTransformSize, - chromaTransformType, - intraLumaDir, - coefficients, - Av1ComponentType.Chroma, - blockContext, - redTransformBlock.EndOfBlock, - frameHeader.UseReducedTransformSet, - blk_ptr.FilterIntraMode, - usesInterTransformSet); - - cb_dc_sign_level_coeff_na.UnitModeWrite( - (byte)culLevelCb, - chromaOrigin, - new Size(transformWidth, transformHeight), - Av1NeighborArrayUnit.UnitMask.Top | Av1NeighborArrayUnit.UnitMask.Left); - - cr_dc_sign_level_coeff_na.UnitModeWrite( - (byte)culLevelCr, - chromaOrigin, - new Size(transformWidth, transformHeight), - Av1NeighborArrayUnit.UnitMask.Top | Av1NeighborArrayUnit.UnitMask.Left); - - entropyCodingContext.CodedAreaSuperblockUv += transformWidth * transformHeight; + for (int blockColumn = regionColumn; blockColumn < unitWidth; blockColumn += transformBlockWidth) + { + int transformStateIndex = + codedArea / Av1EncoderCoefficientBuffer.TransformBlockUnitCoefficientCount; + + ref Av1EncoderTransformBlockState transformBlock = + ref planeTransformBlocks[transformStateIndex]; + + Point chromaOrigin = chromaBlockOrigin + new Size( + blockColumn << Av1Constants.ModeInfoSizeLog2, + blockRow << Av1Constants.ModeInfoSizeLog2); + + Span coefficients = planeCoefficients[codedArea..]; + Av1TransformBlockContext blockContext = GetTransformBlockContexts( + Av1ComponentType.Chroma, + coefficientNeighbors, + chromaOrigin, + chromaBlockSize, + chromaTransformSize); + + int culLevel = writer.WriteCoefficients( + chromaTransformSize, + transformBlock.TransformType, + intraLumaDir, + coefficients, + Av1ComponentType.Chroma, + blockContext, + transformBlock.EndOfBlock, + frameHeader.UseReducedTransformSet, + blk_ptr.FilterIntraMode, + usesInterTransformSet); + + coefficientNeighbors.UnitModeWrite( + (byte)culLevel, + chromaOrigin, + new Size(transformWidth, transformHeight), + Av1NeighborArrayUnit.UnitMask.Top | Av1NeighborArrayUnit.UnitMask.Left); + + codedArea += transformWidth * transformHeight; + } } } } + + codedAreaEnd = codedArea; } + + entropyCodingContext.CodedAreaSuperblockUv = codedAreaEnd; } /// diff --git a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs index 73c6f2d986..d7d8a74125 100644 --- a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs +++ b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs @@ -277,7 +277,7 @@ public class Av1EncoderFrameTests [Fact] public void EncodeEffortTenSelectsThirtyTwoByThirtyTwoBlocks() { - const int Size = 64; + const int Size = 32; using Image source = new(Size, Size); for (int y = 0; y < Size; y++) { @@ -297,6 +297,49 @@ public class Av1EncoderFrameTests qIndex: 4, effort: 10); + byte[] payload = stream.ToArray(); + using Av1Decoder decoder = new(Configuration.Default); + using Image decoded = decoder.Decode(payload); + Av1FrameInfo frameInfo = Assert.IsType(decoder.FrameInfo); + for (int modeInfoY = 0; modeInfoY < 8; modeInfoY++) + { + for (int modeInfoX = 0; modeInfoX < 8; modeInfoX++) + { + Assert.Equal( + Av1BlockSize.Block32x32, + frameInfo.GetModeInfoAt(new Point(modeInfoX, modeInfoY)).BlockSize); + } + } + + Assert.Equal(new Size(Size, Size), decoded.Size); + } + + [Theory] + [InlineData(Yuv400)] + [InlineData(Yuv444)] + public void EncodeEffortTenSelectsSixtyFourBySixtyFourBlock(int colorFormatValue) + { + const int Size = 64; + Av1ColorFormat colorFormat = (Av1ColorFormat)colorFormatValue; + using Image source = new(Size, Size); + for (int y = 0; y < Size; y++) + { + Span row = source.Frames.RootFrame.PixelBuffer.DangerousGetRowSpan(y); + for (int x = 0; x < Size; x++) + { + row[x] = new Rgba32(180, 64, 220); + } + } + + using MemoryStream stream = new(); + _ = Av1FrameEncoder.Encode( + Configuration.Default, + source.Frames.RootFrame, + stream, + CreateColorConfig(Av1BitDepth.EightBit, colorFormat), + qIndex: 4, + effort: 10); + byte[] payload = stream.ToArray(); using Av1Decoder decoder = new(Configuration.Default); using Image decoded = decoder.Decode(payload); @@ -306,7 +349,7 @@ public class Av1EncoderFrameTests for (int modeInfoX = 0; modeInfoX < 16; modeInfoX++) { Assert.Equal( - Av1BlockSize.Block32x32, + Av1BlockSize.Block64x64, frameInfo.GetModeInfoAt(new Point(modeInfoX, modeInfoY)).BlockSize); } } diff --git a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs index 7b2effd840..0eee294fcb 100644 --- a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs +++ b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs @@ -623,10 +623,9 @@ public class Av1TransformBlockEncoderTests Assert.Equal(Av1EncoderModeDecisionWorkspace.MaximumSampleCount, modeWorkspace.GetCandidateReconstruction(1).Length); Assert.Equal(Av1EncoderModeDecisionWorkspace.MaximumSampleCount, modeWorkspace.GetCandidateCoefficients(1).Length); - Assert.Equal( - Av1ChromaFromLumaContext.BufferLine * - Av1EncoderModeDecisionWorkspace.MaximumBlockDimension, - modeWorkspace.ChromaFromLumaSamples.Length); + + // CfL is unavailable above 32x32, so its scratch remains fixed while larger partitions are enabled. + Assert.Equal(Av1ChromaFromLumaContext.BufferLength, modeWorkspace.ChromaFromLumaSamples.Length); Assert.Equal(Av1ChromaFromLumaMath.AlphaCandidateCount, modeWorkspace.GetChromaFromLumaRates(1).Length); Assert.Equal(Av1ChromaFromLumaMath.AlphaCandidateCount, modeWorkspace.GetChromaFromLumaDistortions(1).Length);