diff --git a/HEIF_IMPLEMENTATION_PLAN.md b/HEIF_IMPLEMENTATION_PLAN.md index e40bce1833..ba854d7280 100644 --- a/HEIF_IMPLEMENTATION_PLAN.md +++ b/HEIF_IMPLEMENTATION_PLAN.md @@ -825,13 +825,13 @@ Encoder verification contract: - [~] A production single-tile all-intra writer now walks raster superblocks, analyzes each immediately before entropy coding, reuses one decision workspace and one block workspace, and retains decoder-identical reconstructed references across the tile. Its closed byte and high-bit-depth operators feed the existing superblock boundary without runtime sample-type checks. A byte-exact test compares this composed path with an explicit superblock-then-tile-writer oracle, so producer and writer traversal or coefficient-area drift cannot pass unnoticed. A separate clipped 2x2-superblock regression proves global raster indexing by requiring all four coefficient segments and the bottom-right reconstruction to be populated. Multi-tile ownership and the complete frame/OBU operation remain. - [~] A non-owning encoder-frame view now separates visible conversion regions from coded regions and performs complete left, top, right, bottom, and corner extension across each bordered plane. Current libaom uses 8-sample-aligned coded dimensions, a 32-sample-aligned luma stride with chroma stride derived from it, and a 64-pixel luma border for non-resized all-intra encoding. One operation-ready frame owner now rents the aligned Y, U, and V storage contiguously, exposes non-owning `Buffer2D` plane views, and returns the rent exactly once. A 4K 4:2:0 frame occupies about 13.0 MiB at 8-bit or 26.0 MiB at 10/12-bit; source and reconstruction therefore remain distinct frame owners rather than adding a full-frame copy. The corrected tests use this real ownership path and verify the exact 54 KiB 64x64 4:2:0 rent. The frame-encoder operation now instantiates matching source and reconstruction owners with ordinary `using` lifetimes and converts packed pixels directly into the source owner before extension. - [~] Temporal delimiter, sequence header, frame header, combined-frame tile-group writing, and an internal reduced-still-picture frame operation now exist locally. The remaining required metadata, padding, multi-tile, option, and public encoder paths are not complete. -- [~] Implement superblock and partition analysis for every permitted block size and partition. Efforts zero through eight deliberately split every in-frame node to 8x8 blocks. Effort nine performs recursive live rate-distortion selection at complete 8x8 and 16x16 nodes, while effort ten extends the same search to complete 32x32 and 64x64 nodes. Candidate order matches current libaom: `PARTITION_NONE`, `PARTITION_SPLIT`, `PARTITION_HORZ`, `PARTITION_VERT`, the four asymmetric partitions, then `PARTITION_HORZ_4` and `PARTITION_VERT_4`. Invalid chroma geometries are excluded before evaluation. Each candidate saves and restores the exact partition, coefficient, transform, and palette neighbor edges in one aligned block-workspace owner; trials neither allocate nor copy probability state. Recursive split trials publish each selected child's decoded mode, transform, coefficient, and palette contexts before evaluating its next sibling. Coefficient contexts are published per retained transform rather than broadcasting the first transform over an entire partition leaf. Large chroma leaves are evaluated as raster-ordered transform tiles in the existing aligned workspace, and each winning plane is copied to retained storage once. Production picture state retains the compact 8x8 mode allocation below effort nine and explicitly selects 4x4 allocation granularity when sub-8x8 partitions are enabled. Partition selection above 64x64 and effort-dependent pruning remain. +- [~] Implement superblock and partition analysis for every permitted block size and partition. Efforts zero through eight deliberately split every in-frame node to 8x8 blocks. Effort nine performs recursive live rate-distortion selection at complete 8x8 and 16x16 nodes, while effort ten extends the same search to complete 32x32, 64x64, and 128x128 nodes. Candidate order matches current libaom: `PARTITION_NONE`, `PARTITION_SPLIT`, `PARTITION_HORZ`, `PARTITION_VERT`, the four asymmetric partitions, then `PARTITION_HORZ_4` and `PARTITION_VERT_4`; the two 1-to-4 partitions are excluded at 128x128 as required by current libaom. Invalid chroma geometries are excluded before evaluation. Each candidate saves and restores the exact partition, coefficient, transform, and palette neighbor edges in one aligned block-workspace owner; trials neither allocate nor copy probability state. Recursive split trials publish each selected child's decoded mode, transform, coefficient, and palette contexts before evaluating its next sibling. Coefficient contexts are published per retained transform rather than broadcasting the first transform over an entire partition leaf. Large luma and chroma leaves are evaluated as bounded-64, raster-ordered transform tiles in the existing aligned workspace, and each winning plane is copied to retained storage once. Production picture state retains the compact 8x8 mode allocation below effort nine and explicitly selects 4x4 allocation granularity when sub-8x8 partitions are enabled. Effort-dependent pruning remains. - [~] Implement intra mode search, palette, filter intra, chroma-from-luma, and intra-block copy decisions. Live luma search now covers all 13 zero-angle base modes and all six nonzero adjustments for each of the eight directional modes. Joint spatial chroma search covers the same 61 candidates, combines both chroma planes in one rate-distortion decision, and preserves the winning shared angle adjustment. Chroma-from-luma now searches the complete signed alpha alphabet from reconstructed luma and retains its joint U/V syntax. Filter-intra now searches all five predictors after ordinary luma modes. Palette entropy, retained state, production syntax, exhaustive luma and paired chroma palette selection, adaptive screen-content activation, and joint intra-block-copy mode selection are complete. - [ ] Implement inter mode search for bounded sequences, including reference selection and the decoder-supported inter tools. - [~] Current-libaom `av1_quantize_fp_no_qmatrix` arithmetic is implemented as a closed generic forward-quantizer family with Vector512, Vector256, Vector128, and scalar paths, raster-order output, coded 64-point coefficient limits, and scan-order EOB selection. Transform search, coefficient optimization, and lossless behavior remain. -- [~] Implement real rate-distortion selection and make quality and effort change work, size, and output quality. The complete luma and joint chroma candidate sets, including chroma-from-luma, filter-intra, palette, and intra-block copy, now perform live rate-distortion selection. Public quality mapping and effort tiers through exhaustive uniform luma mode/transform search are implemented. Effort nine adds exact recursive 8x8 and 16x16 partition rate-distortion selection, and effort ten extends it through 64x64; partition search above 64x64, effort-dependent pruning, and the remaining sequence searches remain. -- [~] Frame effort now progressively expands the available current search: zero is DC-only, one adds every zero-angle spatial mode, two adds every legal directional adjustment, three refines the preliminary luma winner's transform type, four adds filter-intra and chroma-from-luma, and five adds adaptive palette and intra-block-copy analysis. Lower tiers do not signal unavailable sequence or frame tools, and tiers below five skip the whole-frame screen-content scan. Effort six enables `TX_MODE_SELECT` and compares the winning ordinary spatial or filter-intra luma mode as one 8x8 transform against four raster-ordered 4x4 transforms; each luma palette candidate owns that size comparison from effort six onward. Effort seven searches every legal 8x8 transform type inside every ordinary spatial candidate rather than refining only the preliminary winner. Effort eight also performs the 8x8-versus-four-4x4 comparison inside every ordinary spatial and filter-intra candidate, matching current libaom's per-candidate uniform-transform ownership. Effort nine additionally searches every legal partition at complete 8x8 and 16x16 nodes in current-libaom order, and effort ten extends that recursive search through 64x64. Prediction and residual construction run once per mode and are reused across its legal transform types. A 64x64 4:4:4 leaf evaluates four 32x32 transforms per chroma plane, retains sparse transform state at coefficient-area offsets, and emits every U transform before every V transform as required by AV1 residual traversal. Every 4x4 transform searches all legal types with live coefficient contexts and reconstructed intra references. The search reuses the aligned block workspace, preserves only global improvements, and performs no per-block, per-partition, or per-transform rent. Non-skipped intra-block copy writes and costs the current-libaom unsplit variable-transform root; skipped intra-block copy emits no transform-partition symbol. Partition search above 64x64 and effort-dependent model/transform pruning remain. Decoder-visible production cases inspect the emitted restrictions and frame state and decode the produced streams, including real effort-nine streams selecting sub-8x8 and 8x16 rectangular blocks. The complete non-HEVC HEIF/AV1 namespace passes 9,074 of 9,074 through one foreground net11 Release VSTest run. The last independently built `aomdec`, from the then-current `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727` snapshot, accepts the previously generated effort-eight and effort-ten payloads as well as the existing palette and intra-block-copy payloads. The affected encoder, partition, and workspace surface passes 139 of 139 through one foreground net11 Release VSTest run. The net11 Release build and Roslynk compiler and analyzer passes report zero errors. -- [~] Encoder rate accounting converts the entropy writer's live inverse cumulative distributions into current-libaom fixed-point symbol costs without allocating or duplicating probability state. Read-only luma-mode, directional-delta, filter-intra, chroma-mode, block-skip, transform-size, transform-block-skip, and complete transform-coefficient queries share the exact distributions mutated by the subsequent entropy write. Complete coefficient costing follows current libaom's optimized shape: it returns immediately for an empty transform, uses the EOB-specific base-range context, fuses magnitude, sign, base-range, and Golomb accounting into one reverse traversal, and combines repeated full base-range chunks instead of replaying each emitted symbol. Tile-lifetime level and context scratch is reused, the one-coefficient path neither clears nor initializes the forward-neighbor level map, and steady-state queries allocate nothing. Transform-size writing and costing share one subdivision-depth calculation, while shared closed symbol operations keep the writer and cost mappings for transform skip, transform type, and EOB syntax identical without forcing the estimator through the writer's slower two-pass coefficient traversal. The current-libaom fixed-point RD combiner preserves 64-bit distortion and rounds the weighted 1/512-bit rate at the required boundary. Its key-frame multiplier follows libaom's squared DC-quantizer formula and exact 10/12-bit normalization. Live final-block selection evaluates all 61 legal 8x8 luma candidates: the 13 zero-angle base modes in current-libaom order, followed by six nonzero adjustments for each directional mode. Joint chroma selection evaluates the equivalent 61 spatial candidates, combines U and V distortion plus coefficient rate, and charges one live chroma-mode and shared-angle symbol over the actual subsampled 4x4, 4x8, or 8x8 geometry. Chroma-from-luma subsamples the reconstructed luma block once into fixed-stride Q3 stack scratch, subtracts the rounded mean, evaluates all 33 signed alpha values independently for each plane with complete transform RD, and combines the cached plane results across all 1,088 valid joint pairs with one live sign cost and the conditional U/V magnitude costs. This is the allocation-free equivalent of current libaom's exhaustive 33-value path: it requires 66 evaluation transforms rather than transforming every joint pair, preserves DC-before-CfL-before-spatial tie order, and fixes the implicit chroma transform to DCT-DCT. Filter-intra follows ordinary luma candidates, searches all five predictors in syntax order, and evaluates every legal transform while reusing one prepared prediction and source residual per filter mode. Every candidate includes its live mode, angle, filter mode, alpha, and coefficient rate plus normalized pixel-domain distortion. Each prepared reference edge retains the common-corner prefix and twice the transform dimension required by directional prediction. A shared encoder/decoder availability calculation selects reconstructed top-right and bottom-left extensions according to tile, frame, superblock, and block reconstruction order; unavailable extensions repeat the nearest coded endpoint. Missing top or left edges retain current libaom's perpendicular-sample and bit-depth-midpoint rules. Directional prediction applies the AV1 three-degree adjustment step and reuses transform workspace for zone-three transposition before the transform overwrites it, keeping candidate evaluation allocation-free. The winning luma and chroma signed adjustments are retained in the packed final-block state consumed by the tile writer. The tile writer invokes these reusable workspace-backed selectors after mapping current neighbors and immediately before writing each block, so later decisions see reconstructed samples, coefficient contexts, and CDF updates from every preceding block. Block skip is read only after the callback has combined every coded plane. Luma and chroma candidate scratch is partitioned from the encoder's single aligned reusable block workspace; transform-size search uses that owner for four retained 4x4 transform states, local coefficient contexts, and the compact trial reconstruction needed to preserve the best result. No candidate path rents a buffer per block or per transform. Only a newly winning candidate is copied into retained frame storage. Production fixtures force every luma base predictor, both extreme adjustments in all three directional zones, available top-right and bottom-left extensions, high-bit-depth adjustment propagation, exact signed luma and chroma angle-rate terms, joint U/V decisions, packed chroma state, and 4:2:0, 4:2:2, and 4:4:4 transform geometry. The CfL fixtures derive target chroma from a pilot production encode's actual reconstructed luma through an independent scalar Q3 oracle and prove exact positive/negative alpha syntax plus zero-residual DCT-DCT reconstruction for all three subsampling geometries at 8, 10, and 12 bits. The stable fixed-DC traversal comparison uses neutral samples for which both the baseline and live search are contractually DC and skipped, instead of relying on textured content to happen to select the baseline mode. Luma palette selection now evaluates dominant-color and one-dimensional K-means candidates for every legal size, snaps near-cache colors with the reference threshold and tie order, removes duplicate snapped colors, extends boundary maps from active samples, and performs complete transform rate-distortion search. Ordinary DC and filter-intra candidates pay the palette-disabled symbol whenever screen-content syntax is enabled. The exact net11 Release rebuild reports 1,992 test-project warnings and zero errors, all 58 intra-superblock cases pass, all 8,935 AVIF cases pass, and all 230 HEIF cases pass. Remaining mode decision work includes transform-size coverage for filter-intra and palette, broader joint mode/transform refinement, partition search above 64x64 and effort-dependent pruning. Non-empty intra blocks deliberately remain non-skipped, matching current libaom; later inter mode selection owns its distinct skip-transform RD decision. +- [~] Implement real rate-distortion selection and make quality and effort change work, size, and output quality. The complete luma and joint chroma candidate sets, including chroma-from-luma, filter-intra, palette, and intra-block copy, now perform live rate-distortion selection. Public quality mapping and effort tiers through exhaustive uniform luma mode/transform search are implemented. Effort nine adds exact recursive 8x8 and 16x16 partition rate-distortion selection, and effort ten extends it through 128x128; effort-dependent pruning and the remaining sequence searches remain. +- [~] Frame effort now progressively expands the available current search: zero is DC-only, one adds every zero-angle spatial mode, two adds every legal directional adjustment, three refines the preliminary luma winner's transform type, four adds filter-intra and chroma-from-luma, and five adds adaptive palette and intra-block-copy analysis. Lower tiers do not signal unavailable sequence or frame tools, and tiers below five skip the whole-frame screen-content scan. Effort six enables `TX_MODE_SELECT` and compares the winning ordinary spatial or filter-intra luma mode as one 8x8 transform against four raster-ordered 4x4 transforms; each luma palette candidate owns that size comparison from effort six onward. Effort seven searches every legal 8x8 transform type inside every ordinary spatial candidate rather than refining only the preliminary winner. Effort eight also performs the 8x8-versus-four-4x4 comparison inside every ordinary spatial and filter-intra candidate, matching current libaom's per-candidate uniform-transform ownership. Effort nine additionally searches every legal partition at complete 8x8 and 16x16 nodes in current-libaom order, and effort ten extends that recursive search through 128x128. Prediction and residual construction run once per mode and are reused across its legal transform types. A 128x128 leaf evaluates four 64x64 luma transforms and as many as sixteen 32x32 transforms per 4:4:4 chroma plane, retaining sparse state at coefficient-area offsets. Residual emission follows AV1's bounded-region order, completing Y, U, and V for each 64x64 luma region before advancing. Every 4x4 transform searches all legal types with live coefficient contexts and reconstructed intra references. The search reuses the aligned block workspace, preserves only global improvements, and performs no per-block, per-partition, or per-transform rent. Non-skipped intra-block copy writes and costs the current-libaom unsplit variable-transform root; skipped intra-block copy emits no transform-partition symbol. Effort-dependent model and transform pruning remain. Decoder-visible production cases inspect the emitted restrictions and frame state and decode the produced streams, including real effort-nine streams selecting sub-8x8 and 8x16 rectangular blocks. The complete non-HEVC HEIF/AV1 namespace passes 9,077 of 9,077 through one foreground net11 Release VSTest run. The last independently built `aomdec`, from the then-current `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727` snapshot, accepts the previously generated effort-eight and effort-ten payloads as well as the existing palette and intra-block-copy payloads. The affected encoder, partition, and workspace surface passes 139 of 139 through one foreground net11 Release VSTest run. The net11 Release build and Roslynk compiler and analyzer passes report zero errors. +- [~] Encoder rate accounting converts the entropy writer's live inverse cumulative distributions into current-libaom fixed-point symbol costs without allocating or duplicating probability state. Read-only luma-mode, directional-delta, filter-intra, chroma-mode, block-skip, transform-size, transform-block-skip, and complete transform-coefficient queries share the exact distributions mutated by the subsequent entropy write. Complete coefficient costing follows current libaom's optimized shape: it returns immediately for an empty transform, uses the EOB-specific base-range context, fuses magnitude, sign, base-range, and Golomb accounting into one reverse traversal, and combines repeated full base-range chunks instead of replaying each emitted symbol. Tile-lifetime level and context scratch is reused, the one-coefficient path neither clears nor initializes the forward-neighbor level map, and steady-state queries allocate nothing. Transform-size writing and costing share one subdivision-depth calculation, while shared closed symbol operations keep the writer and cost mappings for transform skip, transform type, and EOB syntax identical without forcing the estimator through the writer's slower two-pass coefficient traversal. The current-libaom fixed-point RD combiner preserves 64-bit distortion and rounds the weighted 1/512-bit rate at the required boundary. Its key-frame multiplier follows libaom's squared DC-quantizer formula and exact 10/12-bit normalization. Live final-block selection evaluates all 61 legal 8x8 luma candidates: the 13 zero-angle base modes in current-libaom order, followed by six nonzero adjustments for each directional mode. Joint chroma selection evaluates the equivalent 61 spatial candidates, combines U and V distortion plus coefficient rate, and charges one live chroma-mode and shared-angle symbol over the actual subsampled 4x4, 4x8, or 8x8 geometry. Chroma-from-luma subsamples the reconstructed luma block once into fixed-stride Q3 stack scratch, subtracts the rounded mean, evaluates all 33 signed alpha values independently for each plane with complete transform RD, and combines the cached plane results across all 1,088 valid joint pairs with one live sign cost and the conditional U/V magnitude costs. This is the allocation-free equivalent of current libaom's exhaustive 33-value path: it requires 66 evaluation transforms rather than transforming every joint pair, preserves DC-before-CfL-before-spatial tie order, and fixes the implicit chroma transform to DCT-DCT. Filter-intra follows ordinary luma candidates, searches all five predictors in syntax order, and evaluates every legal transform while reusing one prepared prediction and source residual per filter mode. Every candidate includes its live mode, angle, filter mode, alpha, and coefficient rate plus normalized pixel-domain distortion. Each prepared reference edge retains the common-corner prefix and twice the transform dimension required by directional prediction. A shared encoder/decoder availability calculation selects reconstructed top-right and bottom-left extensions according to tile, frame, superblock, and block reconstruction order; unavailable extensions repeat the nearest coded endpoint. Missing top or left edges retain current libaom's perpendicular-sample and bit-depth-midpoint rules. Directional prediction applies the AV1 three-degree adjustment step and reuses transform workspace for zone-three transposition before the transform overwrites it, keeping candidate evaluation allocation-free. The winning luma and chroma signed adjustments are retained in the packed final-block state consumed by the tile writer. The tile writer invokes these reusable workspace-backed selectors after mapping current neighbors and immediately before writing each block, so later decisions see reconstructed samples, coefficient contexts, and CDF updates from every preceding block. Block skip is read only after the callback has combined every coded plane. Luma and chroma candidate scratch is partitioned from the encoder's single aligned reusable block workspace; transform-size search uses that owner for four retained 4x4 transform states, local coefficient contexts, and the compact trial reconstruction needed to preserve the best result. No candidate path rents a buffer per block or per transform. Only a newly winning candidate is copied into retained frame storage. Production fixtures force every luma base predictor, both extreme adjustments in all three directional zones, available top-right and bottom-left extensions, high-bit-depth adjustment propagation, exact signed luma and chroma angle-rate terms, joint U/V decisions, packed chroma state, and 4:2:0, 4:2:2, and 4:4:4 transform geometry. The CfL fixtures derive target chroma from a pilot production encode's actual reconstructed luma through an independent scalar Q3 oracle and prove exact positive/negative alpha syntax plus zero-residual DCT-DCT reconstruction for all three subsampling geometries at 8, 10, and 12 bits. The stable fixed-DC traversal comparison uses neutral samples for which both the baseline and live search are contractually DC and skipped, instead of relying on textured content to happen to select the baseline mode. Luma palette selection now evaluates dominant-color and one-dimensional K-means candidates for every legal size, snaps near-cache colors with the reference threshold and tie order, removes duplicate snapped colors, extends boundary maps from active samples, and performs complete transform rate-distortion search. Ordinary DC and filter-intra candidates pay the palette-disabled symbol whenever screen-content syntax is enabled. The exact net11 Release rebuild reports 1,992 test-project warnings and zero errors, all 58 intra-superblock cases pass, all 8,935 AVIF cases pass, and all 230 HEIF cases pass. Remaining mode decision work includes transform-size coverage for filter-intra and palette, broader joint mode/transform refinement, and effort-dependent pruning. Non-empty intra blocks deliberately remain non-skipped, matching current libaom; later inter mode selection owns its distinct skip-transform RD decision. - [~] The tile writer now publishes one packed coefficient context per covered 4x4 edge unit and derives luma/chroma skip plus DC-sign contexts from the complete transform edges using current-libaom units. Partition, transform, and coefficient neighbor state retains only the above and left context regions used by current libaom; the unused third top-left region, its granularity state, and its unused sentinel are removed. One picture owner now packs segmentation plus every tile's partition, luma, chroma, and transform edges into one clean byte allocation with typed non-owning views; together with the separately typed packed mode-information owner, the complete picture state uses two allocator rents rather than seven. Exact aligned lengths, clean initialization, and balanced exactly-once returns are covered in Release. Multi-tile payload ownership and verified CDF update behavior remain. - [~] Encoder mode information now uses a frame-owned integer alias grid over a packed 8-byte value allocation, matching current libaom's `mi_grid_base` and `mi_alloc` relationship without a managed object or reference per 4x4 entry. The visible dimensions are aligned to eight luma samples, the grid stride and allocated row count are aligned to 32 mode-information units, and optional 8x8 allocation granularity reduces the value store in both dimensions exactly as current libaom does. One clean ImageSharp byte owner contains both independently typed regions, reducing libaom's two allocation lifetimes to one without a copy. At 4K, the 4x4 layout occupies about 6.0 MiB in total; the 8x8 layout occupies about 3.0 MiB. Exact geometry, clean allocation, typed lengths, aligned mapping, untouched row padding, and exactly-once return pass 4 of 4 direct net11 VSTest cases in Release. Every coded 4x4 cell covered by square, rectangular, or clipped edge blocks maps to its owning allocation entry before context-dependent symbols are written. Packed syntax, relative neighbor lookup, full block mapping, writer traversal, entropy, and OBU coverage pass 1,947 of 1,947 direct net11 VSTest cases in Release; complete mode decision still remains. - [~] The final-block decision workspace uses one reusable 8.3 KiB ImageSharp allocator owner. It contains 1,024 explicitly packed 8-byte final-block entries and the 341 preorder partition bytes required by a complete 128x128-through-8x8 quadtree, replacing separate managed arrays. Palette colors now have their own current-block value and are copied only to the picture edges that later blocks can reference, so enabling palette mode does not add 50 bytes to every final-block entry. Construction and the explicit per-superblock reset initialize every syntax field, including the nonzero sentinel that disables filter-intra prediction; pooled quantizer, prediction, partition, and current-palette bytes cannot leak into the next decision pass. Complete mode decision still remains. diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs index c2664d4052..2eedd1837c 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs @@ -18,29 +18,35 @@ internal readonly ref struct Av1EncoderModeDecisionWorkspace /// /// The largest coding-block dimension evaluated directly by the current partition search. /// - public const int MaximumBlockDimension = 64; + public const int MaximumBlockDimension = 128; /// /// The maximum number of samples in one directly evaluated coding block. /// public const int MaximumSampleCount = MaximumBlockDimension * MaximumBlockDimension; + /// + /// The maximum number of samples in one AV1 transform. + /// + public const int MaximumTransformSampleCount = + Av1Constants.MaxTransformSize * Av1Constants.MaxTransformSize; + /// /// The number of 4x4 transform blocks covering one 8x8 coding block. /// public const int CandidateTransformBlockCount = 4; /// - /// The maximum number of transform states needed while evaluating both chroma planes of one 64x64 block. + /// The maximum number of transform states needed while evaluating both chroma planes of one 128x128 block. /// - public const int MaximumCandidateTransformBlockCount = 8; + public const int MaximumCandidateTransformBlockCount = 32; /// /// The required workspace length in signed-integer storage elements. /// public const int StorageLength = TransientStorageOffset + Av1EncoderPaletteWorkspace.StorageLength; - private const int ReferenceBufferLength = (2 * MaximumBlockDimension) + 1; + private const int ReferenceBufferLength = (2 * Av1Constants.MaxTransformSize) + 1; private const int ReferenceBufferCount = 4; private const int ReferenceStorageLength = ReferenceBufferCount * ReferenceBufferLength * sizeof(ushort) / sizeof(int); private const int CandidateSampleStorageOffset = ReferenceStorageLength; @@ -80,7 +86,7 @@ internal readonly ref struct Av1EncoderModeDecisionWorkspace /// Gets the temporary prediction span shared by mutually exclusive mode searches. /// public Span Prediction - => MemoryMarshal.Cast(this.storage[TransientStorageOffset..])[..MaximumSampleCount]; + => MemoryMarshal.Cast(this.storage[TransientStorageOffset..])[..MaximumTransformSampleCount]; /// /// Gets the temporary residual span shared by mutually exclusive mode searches. @@ -88,8 +94,8 @@ internal readonly ref struct Av1EncoderModeDecisionWorkspace public Span Residual => MemoryMarshal.Cast( this.storage.Slice( - TransientStorageOffset + (MaximumSampleCount * sizeof(ushort) / sizeof(int)), - MaximumSampleCount * sizeof(short) / sizeof(int))); + TransientStorageOffset + (MaximumTransformSampleCount * sizeof(ushort) / sizeof(int)), + MaximumTransformSampleCount * sizeof(short) / sizeof(int))); /// /// Gets the fixed-stride subsampled luma values used by chroma-from-luma mode search. @@ -182,7 +188,8 @@ internal readonly ref struct Av1EncoderPaletteWorkspace /// public const int StorageLength = ColorCacheOffset + ColorCacheStorageLength; - private const int MaximumSampleCount = Av1EncoderModeDecisionWorkspace.MaximumSampleCount; + private const int MaximumBlockDimension = 64; + private const int MaximumSampleCount = MaximumBlockDimension * MaximumBlockDimension; private const int PlaneShortStorageLength = MaximumSampleCount * sizeof(short) / sizeof(int); private const int PlaneSampleStorageLength = MaximumSampleCount * sizeof(ushort) / sizeof(int); private const int PlaneByteStorageLength = MaximumSampleCount / sizeof(int); diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1FrameEncoder.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1FrameEncoder.cs index 880d59dcdf..58769c65c2 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1FrameEncoder.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1FrameEncoder.cs @@ -89,7 +89,7 @@ internal static class Av1FrameEncoder FrameHeightBits = height > 1 ? Av1Math.MostSignificantBit((uint)(height - 1)) + 1 : 1, MaxFrameWidth = width, MaxFrameHeight = height, - Use128x128Superblock = false, + Use128x128Superblock = effort == 10 && width >= 128 && height >= 128, ForceScreenContentTools = 2, ForceIntegerMotionVector = 2, EnableFilterIntra = effort >= 4, diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs index e454da6e36..13d32cb518 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs @@ -217,8 +217,9 @@ internal static partial class Av1IntraSuperblockEncoder }; bool hasLumaPalette = paletteInfo.PaletteSizes[0] != 0; - int paletteDisabledCost = blockSize >= Av1BlockSize.Block8x8 && - this.picture.Parent.FrameHeader.AllowScreenContentTools + int paletteDisabledCost = Av1TileWriter.IsPaletteAllowed( + this.picture.Parent.FrameHeader.AllowScreenContentTools, + blockSize) ? writer.GetPaletteUvModeCost(false, hasLumaPalette) : 0; @@ -575,6 +576,9 @@ internal static partial class Av1IntraSuperblockEncoder int transformColumnCount = blockWidth / transformWidth; int transformRowCount = blockHeight / transformHeight; int transformBlockCount = transformColumnCount * transformRowCount; + Av1BlockSize maximumUnitBlockSize = + Av1BlockSize.Block64x64.GetSubsampled(colorConfig.SubSamplingX, colorConfig.SubSamplingY); + Span candidateBlueReconstruction = workspace.GetCandidateReconstruction(0)[..blockSampleCount]; @@ -619,7 +623,9 @@ internal static partial class Av1IntraSuperblockEncoder Buffer2DRegion blueReconstruction = this.reconstruction.GetPlane(Av1Plane.U); Buffer2DRegion redReconstruction = this.reconstruction.GetPlane(Av1Plane.V); bool hasLumaPalette = paletteInfo.PaletteSizes[0] != 0; - int paletteDisabledCost = this.picture.Parent.FrameHeader.AllowScreenContentTools + int paletteDisabledCost = Av1TileWriter.IsPaletteAllowed( + this.picture.Parent.FrameHeader.AllowScreenContentTools, + blockSize) ? writer.GetPaletteUvModeCost(false, hasLumaPalette) : 0; @@ -662,7 +668,7 @@ internal static partial class Av1IntraSuperblockEncoder redNeighbors.Top.Slice(redTopIndex, contextWidth).CopyTo(redTopContexts); redNeighbors.Left.Slice(redLeftIndex, contextHeight).CopyTo(redLeftContexts); Av1PredictionMode predictionMode = chromaMode.ToLumaMode(); - long distortion = this.GetTiledChromaPlaneCost( + long distortion = this.GetTiledPlaneCost( writer, macroBlock, lumaOrigin, @@ -670,6 +676,7 @@ internal static partial class Av1IntraSuperblockEncoder blockSize, chromaBlockSize, transformSize, + maximumUnitBlockSize, subsamplingX, subsamplingY, lumaMode, @@ -685,7 +692,7 @@ internal static partial class Av1IntraSuperblockEncoder blueLeftContexts, out int blueRate); - distortion += this.GetTiledChromaPlaneCost( + distortion += this.GetTiledPlaneCost( writer, macroBlock, lumaOrigin, @@ -693,6 +700,7 @@ internal static partial class Av1IntraSuperblockEncoder blockSize, chromaBlockSize, transformSize, + maximumUnitBlockSize, subsamplingX, subsamplingY, lumaMode, @@ -761,7 +769,7 @@ internal static partial class Av1IntraSuperblockEncoder return bestMode; } - private long GetTiledChromaPlaneCost( + private long GetTiledPlaneCost( Av1SymbolEncoder writer, Av1MacroBlockD macroBlock, Point lumaOrigin, @@ -769,6 +777,7 @@ internal static partial class Av1IntraSuperblockEncoder Av1BlockSize blockSize, Av1BlockSize chromaBlockSize, Av1TransformSize transformSize, + Av1BlockSize maximumUnitBlockSize, int subsamplingX, int subsamplingY, Av1PredictionMode lumaMode, @@ -794,12 +803,17 @@ internal static partial class Av1IntraSuperblockEncoder int transformSampleCount = transformSize.GetSize2d(); int transformWidth4x4 = transformSize.Get4x4WideCount(); int transformHeight4x4 = transformSize.Get4x4HighCount(); + int maximumUnitWidth = Math.Min(maximumUnitBlockSize.GetWidth(), blockWidth); + int maximumUnitHeight = Math.Min(maximumUnitBlockSize.GetHeight(), blockHeight); Av1TransformType transformType = Av1SymbolContextHelper.GetDefaultIntraTransformType( predictionMode, transformSize, this.picture.Parent.FrameHeader.UseReducedTransformSet); - Av1ComponentType componentType = Av1ComponentType.Chroma; + Av1ComponentType componentType = plane == Av1Plane.Y + ? Av1ComponentType.Luminance + : Av1ComponentType.Chroma; + Span prediction = workspace.Prediction[..transformSampleCount]; Span residual = workspace.Residual[..transformSampleCount]; Span aboveStorage = workspace.GetReferenceSamples(0); @@ -808,104 +822,115 @@ internal static partial class Av1IntraSuperblockEncoder int transformIndex = 0; long distortion = 0; rate = 0; - for (int transformRow = 0; transformRow < blockHeight / transformHeight; transformRow++) + + // Residual syntax completes each bounded 64x64 luma region, scaled for chroma, before + // moving to the next region. Candidate coefficients and states must retain that exact order. + for (int regionRow = 0; regionRow < blockHeight; regionRow += maximumUnitHeight) { - int rowOffset = transformRow * transformHeight; - for (int transformColumn = 0; transformColumn < blockWidth / transformWidth; transformColumn++) + int unitBottom = Math.Min(regionRow + maximumUnitHeight, blockHeight); + for (int regionColumn = 0; regionColumn < blockWidth; regionColumn += maximumUnitWidth) { - int columnOffset = transformColumn * transformWidth; - int reconstructionOffset = (rowOffset * blockWidth) + columnOffset; - Point transformOrigin = chromaOrigin + new Size(columnOffset, rowOffset); - this.PrepareTransformReferenceSamples( - reconstruction, - lumaOrigin, - chromaOrigin, - blockSize, - macroBlock, - transformRow, - transformColumn, - blockWidth, - transformSize, - subsamplingX, - subsamplingY, - candidateReconstruction, - aboveStorage, - leftStorage, - out bool hasLeft, - out bool hasAbove); - - TOperator.PrepareIntra( - this.blockWorkspace, - source, - transformOrigin, - prediction, - aboveStorage.Slice(1, transformWidth * 2), - leftStorage.Slice(1, transformHeight * 2), - hasLeft, - hasAbove, - predictionMode, - angleDelta, - residual, - transformSize, - this.bitDepth); - - Av1TransformBlockContext blockContext = Av1TileWriter.GetTransformBlockContexts( - componentType, - topContexts.Slice(transformColumn * transformWidth4x4, transformWidth4x4), - leftContexts.Slice(transformRow * transformHeight4x4, transformHeight4x4), - chromaBlockSize, - transformSize); - - Span transformCoefficients = candidateCoefficients.Slice( - coefficientOffset, - transformSampleCount); - - ref Av1EncoderTransformBlockState state = ref candidateStates[transformIndex++]; - distortion += TOperator.EncodePredictionCandidate( - this.blockWorkspace, - source, - transformOrigin, - prediction, - residual, - candidateReconstruction[reconstructionOffset..], - blockWidth, - transformCoefficients, - transformSize, - transformType, - plane, - this.quantization.QIndex[0], - this.quantization.DeltaQDc[(int)plane], - this.quantization.DeltaQAc[(int)plane], - this.bitDepth, - ref state); - - rate += writer.GetCoefficientCost( - transformSize, - transformType, - lumaMode, - transformCoefficients, - componentType, - blockContext, - state.EndOfBlock, - this.picture.Parent.FrameHeader.UseReducedTransformSet, - Av1FilterIntraMode.AllFilterIntraModes, - usesInterTransformSet: false); - - byte coefficientContext = Av1SymbolContextHelper.GetCoefficientContext( - transformCoefficients, - transformSize, - transformType, - state.EndOfBlock); - - topContexts - .Slice(transformColumn * transformWidth4x4, transformWidth4x4) - .Fill(coefficientContext); - - leftContexts - .Slice(transformRow * transformHeight4x4, transformHeight4x4) - .Fill(coefficientContext); - - coefficientOffset += transformSampleCount; + int unitRight = Math.Min(regionColumn + maximumUnitWidth, blockWidth); + for (int rowOffset = regionRow; rowOffset < unitBottom; rowOffset += transformHeight) + { + int transformRow = rowOffset / transformHeight; + for (int columnOffset = regionColumn; columnOffset < unitRight; columnOffset += transformWidth) + { + int transformColumn = columnOffset / transformWidth; + int reconstructionOffset = (rowOffset * blockWidth) + columnOffset; + Point transformOrigin = chromaOrigin + new Size(columnOffset, rowOffset); + this.PrepareTransformReferenceSamples( + reconstruction, + lumaOrigin, + chromaOrigin, + blockSize, + macroBlock, + transformRow, + transformColumn, + blockWidth, + transformSize, + subsamplingX, + subsamplingY, + candidateReconstruction, + aboveStorage, + leftStorage, + out bool hasLeft, + out bool hasAbove); + + TOperator.PrepareIntra( + this.blockWorkspace, + source, + transformOrigin, + prediction, + aboveStorage.Slice(1, transformWidth * 2), + leftStorage.Slice(1, transformHeight * 2), + hasLeft, + hasAbove, + predictionMode, + angleDelta, + residual, + transformSize, + this.bitDepth); + + Av1TransformBlockContext blockContext = Av1TileWriter.GetTransformBlockContexts( + componentType, + topContexts.Slice(transformColumn * transformWidth4x4, transformWidth4x4), + leftContexts.Slice(transformRow * transformHeight4x4, transformHeight4x4), + chromaBlockSize, + transformSize); + + Span transformCoefficients = candidateCoefficients.Slice( + coefficientOffset, + transformSampleCount); + + ref Av1EncoderTransformBlockState state = ref candidateStates[transformIndex++]; + distortion += TOperator.EncodePredictionCandidate( + this.blockWorkspace, + source, + transformOrigin, + prediction, + residual, + candidateReconstruction[reconstructionOffset..], + blockWidth, + transformCoefficients, + transformSize, + transformType, + plane, + this.quantization.QIndex[0], + this.quantization.DeltaQDc[(int)plane], + this.quantization.DeltaQAc[(int)plane], + this.bitDepth, + ref state); + + rate += writer.GetCoefficientCost( + transformSize, + transformType, + lumaMode, + transformCoefficients, + componentType, + blockContext, + state.EndOfBlock, + this.picture.Parent.FrameHeader.UseReducedTransformSet, + Av1FilterIntraMode.AllFilterIntraModes, + usesInterTransformSet: false); + + byte coefficientContext = Av1SymbolContextHelper.GetCoefficientContext( + transformCoefficients, + transformSize, + transformType, + state.EndOfBlock); + + topContexts + .Slice(transformColumn * transformWidth4x4, transformWidth4x4) + .Fill(coefficientContext); + + leftContexts + .Slice(transformRow * transformHeight4x4, transformHeight4x4) + .Fill(coefficientContext); + + coefficientOffset += transformSampleCount; + } + } } } diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs index 5f139a164f..aad34f945d 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs @@ -184,7 +184,8 @@ internal static partial class Av1IntraSuperblockEncoder Av1PartitionType preparedPartition) { bool searchPartition = blockSize is Av1BlockSize.Block8x8 or Av1BlockSize.Block16x16 || - (this.effort == 10 && blockSize is Av1BlockSize.Block32x32 or Av1BlockSize.Block64x64); + (this.effort == 10 && + blockSize is Av1BlockSize.Block32x32 or Av1BlockSize.Block64x64 or Av1BlockSize.Block128x128); if (!searchPartition) { @@ -406,6 +407,13 @@ internal static partial class Av1IntraSuperblockEncoder return false; } + if (blockSize == Av1BlockSize.Block128x128 && + partitionType is Av1PartitionType.Horizontal4 or Av1PartitionType.Vertical4) + { + // AV1 excludes 128x32 and 32x128 leaves from the 128x128 partition alphabet. + return false; + } + if (this.source.IsMonochrome) { return true; @@ -584,13 +592,17 @@ internal static partial class Av1IntraSuperblockEncoder // Mode decision retains one state for every uniform transform tile in the coding block. The block // can skip coefficient syntax only when every retained transform has an empty end-of-block marker. + int lumaTransformSampleCount = lumaTransformSize.GetSize2d(); int lumaTransformBlockCount = - (blockSize.GetWidth() * blockSize.GetHeight()) / lumaTransformSize.GetSize2d(); + (blockSize.GetWidth() * blockSize.GetHeight()) / lumaTransformSampleCount; + + int lumaStateStride = + lumaTransformSampleCount / Av1EncoderCoefficientBuffer.TransformBlockUnitCoefficientCount; bool lumaTransformEmpty = true; for (int transformIndex = 0; transformIndex < lumaTransformBlockCount; transformIndex++) { - lumaTransformEmpty &= retainedLumaStates[transformIndex].EndOfBlock == 0; + lumaTransformEmpty &= retainedLumaStates[transformIndex * lumaStateStride].EndOfBlock == 0; } if (this.source.IsMonochrome) @@ -886,6 +898,7 @@ internal static partial class Av1IntraSuperblockEncoder blockOrigin, blockSize, transformSize, + Av1BlockSize.Block64x64, lumaCoefficients[lumaArea..], lumaStates[(lumaArea / Av1EncoderCoefficientBuffer.TransformBlockUnitCoefficientCount)..]); @@ -923,6 +936,9 @@ internal static partial class Av1IntraSuperblockEncoder colorConfig.SubSamplingX, colorConfig.SubSamplingY); + Av1BlockSize maximumChromaUnitBlockSize = + Av1BlockSize.Block64x64.GetSubsampled(colorConfig.SubSamplingX, colorConfig.SubSamplingY); + int chromaStateIndex = chromaArea / Av1EncoderCoefficientBuffer.TransformBlockUnitCoefficientCount; @@ -939,6 +955,7 @@ internal static partial class Av1IntraSuperblockEncoder chromaOrigin, chromaBlockSize, chromaTransformSize, + maximumChromaUnitBlockSize, blueCoefficients[chromaArea..], blueStates[chromaStateIndex..]); @@ -947,6 +964,7 @@ internal static partial class Av1IntraSuperblockEncoder chromaOrigin, chromaBlockSize, chromaTransformSize, + maximumChromaUnitBlockSize, redCoefficients[chromaArea..], redStates[chromaStateIndex..]); } @@ -956,6 +974,7 @@ internal static partial class Av1IntraSuperblockEncoder Point blockOrigin, Av1BlockSize blockSize, Av1TransformSize transformSize, + Av1BlockSize maximumUnitBlockSize, ReadOnlySpan coefficients, ReadOnlySpan states) { @@ -968,32 +987,42 @@ internal static partial class Av1IntraSuperblockEncoder int transformWidth = transformSize.GetWidth(); int transformHeight = transformSize.GetHeight(); int transformSampleCount = transformSize.GetSize2d(); + int maximumUnitWidth = Math.Min(maximumUnitBlockSize.GetWidth(), blockWidth); + int maximumUnitHeight = Math.Min(maximumUnitBlockSize.GetHeight(), blockHeight); int transformStateOffset = 0; int coefficientOffset = 0; int transformStateStride = transformSampleCount / Av1EncoderCoefficientBuffer.TransformBlockUnitCoefficientCount; - // Uniform transform blocks are retained and written in raster order. Publishing that same tiling - // preserves the distinct top and left contexts consumed by the next coding block in a dry run. - for (int row = 0; row < blockHeight; row += transformHeight) + // Coefficients retain AV1's bounded-region order rather than unrestricted row-major order. + // Publishing the same sequence pairs every state with the transform that produced it. + for (int regionRow = 0; regionRow < blockHeight; regionRow += maximumUnitHeight) { - for (int column = 0; column < blockWidth; column += transformWidth) + int unitBottom = Math.Min(regionRow + maximumUnitHeight, blockHeight); + for (int regionColumn = 0; regionColumn < blockWidth; regionColumn += maximumUnitWidth) { - Av1EncoderTransformBlockState state = states[transformStateOffset]; - byte context = Av1SymbolContextHelper.GetCoefficientContext( - coefficients[coefficientOffset..], - transformSize, - state.TransformType, - state.EndOfBlock); + int unitRight = Math.Min(regionColumn + maximumUnitWidth, blockWidth); + for (int row = regionRow; row < unitBottom; row += transformHeight) + { + for (int column = regionColumn; column < unitRight; column += transformWidth) + { + Av1EncoderTransformBlockState state = states[transformStateOffset]; + byte context = Av1SymbolContextHelper.GetCoefficientContext( + coefficients[coefficientOffset..], + transformSize, + state.TransformType, + state.EndOfBlock); - neighbors.UnitModeWrite( - context, - blockOrigin + new Size(column, row), - new Size(transformWidth, transformHeight), - EdgeMask); + neighbors.UnitModeWrite( + context, + blockOrigin + new Size(column, row), + new Size(transformWidth, transformHeight), + EdgeMask); - coefficientOffset += transformSampleCount; - transformStateOffset += transformStateStride; + coefficientOffset += transformSampleCount; + transformStateOffset += transformStateStride; + } + } } } } @@ -1315,6 +1344,23 @@ internal static partial class Av1IntraSuperblockEncoder Buffer2DRegion sourcePlane = this.source.GetPlane(Av1Plane.Y); Buffer2DRegion reconstructionPlane = this.reconstruction.GetPlane(Av1Plane.Y); + if (blockWidth > transformSize.GetWidth() || blockHeight > transformSize.GetHeight()) + { + return this.SelectTiledLumaMode( + writer, + macroBlock, + blockOrigin, + blockSize, + tileIndex, + transformSize, + retainedCoefficients, + retainedStates, + out selectedAngleDelta, + out selectedFilterIntraMode, + out selectedTransformSize, + out selectedCost); + } + bool hasLeft = macroBlock.IsLeftAvailable; bool hasAbove = macroBlock.IsUpAvailable; int modeInfoRow = blockOrigin.Y >> Av1Constants.ModeInfoSizeLog2; @@ -1438,8 +1484,9 @@ internal static partial class Av1IntraSuperblockEncoder : 0; int paletteDisabledCost = 0; - if (blockSize >= Av1BlockSize.Block8x8 && - this.picture.Parent.FrameHeader.AllowScreenContentTools) + if (Av1TileWriter.IsPaletteAllowed( + this.picture.Parent.FrameHeader.AllowScreenContentTools, + blockSize)) { Av1NeighborArrayUnit paletteContexts = this.picture.PaletteContexts[tileIndex]; int blockSizeContext = Av1TileWriter.GetPaletteBlockSizeContext(blockSize); @@ -1698,9 +1745,9 @@ internal static partial class Av1IntraSuperblockEncoder } if (this.effort >= 4 && - this.picture.Sequence.SequenceHeader.EnableFilterIntra && - blockWidth <= 32 && - blockHeight <= 32) + Av1TileWriter.IsFilterIntraAllowedBlockSize( + this.picture.Sequence.SequenceHeader.EnableFilterIntra, + blockSize)) { // Each recursive filter prediction and its source residual are independent of transform type. // Prepare them once per filter mode so all legal transforms reuse the same samples. @@ -1889,6 +1936,174 @@ internal static partial class Av1IntraSuperblockEncoder return bestMode; } + private Av1PredictionMode SelectTiledLumaMode( + Av1SymbolEncoder writer, + Av1MacroBlockD macroBlock, + Point blockOrigin, + Av1BlockSize blockSize, + ushort tileIndex, + Av1TransformSize transformSize, + Span retainedCoefficients, + Span retainedStates, + out int selectedAngleDelta, + out Av1FilterIntraMode selectedFilterIntraMode, + out Av1TransformSize selectedTransformSize, + out long selectedCost) + { + Av1EncoderModeDecisionWorkspace workspace = + this.blockWorkspace.GetModeDecisionWorkspace(); + + int blockWidth = blockSize.GetWidth(); + int blockHeight = blockSize.GetHeight(); + int blockSampleCount = blockWidth * blockHeight; + int transformBlockCount = blockSampleCount / transformSize.GetSize2d(); + Span candidateReconstruction = + workspace.GetCandidateReconstruction(0)[..blockSampleCount]; + + Span candidateCoefficients = + workspace.GetCandidateCoefficients(0)[..blockSampleCount]; + + Span candidateStates = + workspace.CandidateTransformBlocks[..transformBlockCount]; + + int contextWidth = blockSize.Get4x4WideCount(); + int contextHeight = blockSize.Get4x4HighCount(); + Span contexts = workspace.TransformContexts; + Span topContexts = contexts[..contextWidth]; + Span leftContexts = contexts.Slice(contextWidth, contextHeight); + Av1NeighborArrayUnit coefficientNeighbors = + this.picture.LuminanceDcSignLevelCoefficientNeighbors[tileIndex]; + + int topIndex = coefficientNeighbors.GetTopIndex(blockOrigin); + int leftIndex = coefficientNeighbors.GetLeftIndex(blockOrigin); + int transformSizeContext = Av1TileWriter.GetTransformSizeContext( + this.picture.TransformFunctionContexts[tileIndex], + macroBlock, + blockOrigin, + blockSize); + + int transformSizeRate = this.picture.Parent.FrameHeader.TransformMode == Av1TransformMode.Select + ? writer.GetTransformSizeCost(blockSize, transformSize, transformSizeContext) + : 0; + + int paletteDisabledCost = Av1TileWriter.IsPaletteAllowed( + this.picture.Parent.FrameHeader.AllowScreenContentTools, + blockSize) + ? writer.GetPaletteYModeCost( + false, + Av1TileWriter.GetPaletteBlockSizeContext(blockSize), + Av1TileWriter.GetPaletteYModeContext( + this.picture.PaletteContexts[tileIndex], + macroBlock, + blockOrigin)) + : 0; + + int baseModeCount = LumaModeSearchOrder.Length; + int deltaCount = AngleDeltaSearchOrder.Length; + int directionalModeCount = + (int)Av1PredictionMode.Directional67Degrees - (int)Av1PredictionMode.Vertical + 1; + + int candidateCount = this.effort switch + { + 0 => 1, + 1 => baseModeCount, + _ => baseModeCount + (directionalModeCount * deltaCount) + }; + + Buffer2DRegion sourcePlane = this.source.GetPlane(Av1Plane.Y); + Buffer2DRegion reconstructionPlane = this.reconstruction.GetPlane(Av1Plane.Y); + long bestCost = long.MaxValue; + Av1PredictionMode bestMode = Av1PredictionMode.DC; + selectedAngleDelta = 0; + selectedFilterIntraMode = Av1FilterIntraMode.AllFilterIntraModes; + selectedTransformSize = transformSize; + + // Each candidate starts from the live block-edge contexts. Transform updates remain local until + // that candidate wins, so later modes never inherit state from an earlier trial. + for (int candidateIndex = 0; candidateIndex < candidateCount; candidateIndex++) + { + Av1PredictionMode mode; + int angleDelta; + if (candidateIndex < baseModeCount) + { + mode = LumaModeSearchOrder[candidateIndex]; + angleDelta = 0; + } + else + { + int adjustedIndex = candidateIndex - baseModeCount; + mode = (Av1PredictionMode)( + (int)Av1PredictionMode.Vertical + (adjustedIndex / deltaCount)); + + angleDelta = AngleDeltaSearchOrder[adjustedIndex % deltaCount]; + } + + coefficientNeighbors.Top.Slice(topIndex, contextWidth).CopyTo(topContexts); + coefficientNeighbors.Left.Slice(leftIndex, contextHeight).CopyTo(leftContexts); + long distortion = this.GetTiledPlaneCost( + writer, + macroBlock, + blockOrigin, + blockOrigin, + blockSize, + blockSize, + transformSize, + Av1BlockSize.Block64x64, + 0, + 0, + mode, + mode, + angleDelta, + Av1Plane.Y, + sourcePlane, + reconstructionPlane, + candidateReconstruction, + candidateCoefficients, + candidateStates, + topContexts, + leftContexts, + out int coefficientRate); + + int rate = Av1TileWriter.GetLumaModeCost(writer, macroBlock, blockSize, mode, angleDelta); + rate += transformSizeRate + coefficientRate; + if (mode == Av1PredictionMode.DC) + { + rate += paletteDisabledCost; + if (Av1TileWriter.IsFilterIntraAllowedBlockSize( + this.picture.Sequence.SequenceHeader.EnableFilterIntra, + blockSize)) + { + rate += writer.GetFilterIntraModeCost( + Av1FilterIntraMode.AllFilterIntraModes, + blockSize); + } + } + + long candidateCost = Av1RateDistortion.GetCost(this.rateMultiplier, rate, distortion); + if (candidateCost < bestCost) + { + CopyTiledCandidate( + candidateReconstruction, + candidateCoefficients, + candidateStates, + reconstructionPlane, + blockOrigin, + blockWidth, + blockHeight, + transformSize, + retainedCoefficients, + retainedStates); + + bestCost = candidateCost; + bestMode = mode; + selectedAngleDelta = angleDelta; + } + } + + selectedCost = bestCost; + return bestMode; + } + private long GetSplitLumaCandidateCost( Av1SymbolEncoder writer, Av1MacroBlockD macroBlock, @@ -2330,10 +2545,21 @@ internal static partial class Av1IntraSuperblockEncoder if (hasBottomLeft) { - for (int row = transformHeight; row < transformHeight * 2; row++) + if (transformColumn > 0) + { + for (int row = transformHeight; row < transformHeight * 2; row++) + { + left[row] = candidateReconstruction[ + ((rowOffset + row) * planeBlockWidth) + columnOffset - 1]; + } + } + else { - left[row] = reconstructionPlane - .DangerousGetRowSpan(planeBlockOrigin.Y + rowOffset + row)[planeBlockOrigin.X - 1]; + for (int row = transformHeight; row < transformHeight * 2; row++) + { + left[row] = reconstructionPlane + .DangerousGetRowSpan(planeBlockOrigin.Y + rowOffset + row)[planeBlockOrigin.X - 1]; + } } } else @@ -2406,7 +2632,10 @@ internal static partial class Av1IntraSuperblockEncoder rate += paletteDisabledCost; } - if (mode == Av1PredictionMode.DC && this.picture.Sequence.SequenceHeader.EnableFilterIntra) + if (mode == Av1PredictionMode.DC && + Av1TileWriter.IsFilterIntraAllowedBlockSize( + this.picture.Sequence.SequenceHeader.EnableFilterIntra, + blockSize)) { rate += writer.GetFilterIntraModeCost(Av1FilterIntraMode.AllFilterIntraModes, blockSize); } diff --git a/src/ImageSharp/Formats/Heif/Av1/Tiling/Av1TileWriter.cs b/src/ImageSharp/Formats/Heif/Av1/Tiling/Av1TileWriter.cs index 45b4b7a3bf..53b591ea4b 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Tiling/Av1TileWriter.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Tiling/Av1TileWriter.cs @@ -1452,7 +1452,7 @@ internal partial class Av1TileWriter /// A value indicating whether the sequence enables filter-intra prediction. /// The block size. /// when filter-intra prediction supports the block dimensions; otherwise, . - private static bool IsFilterIntraAllowedBlockSize(bool enableFilterIntra, Av1BlockSize blockSize) + internal static bool IsFilterIntraAllowedBlockSize(bool enableFilterIntra, Av1BlockSize blockSize) { if (!enableFilterIntra) { @@ -1573,7 +1573,7 @@ internal partial class Av1TileWriter /// A value indicating whether screen-content tools are enabled. /// The block size. /// when palette mode is available for the block; otherwise, . - private static bool IsPaletteAllowed(bool allowScreenContentTools, Av1BlockSize blockSize) + internal static bool IsPaletteAllowed(bool allowScreenContentTools, Av1BlockSize blockSize) => allowScreenContentTools && blockSize.GetWidth() <= 64 && blockSize.GetHeight() <= 64 && @@ -1700,19 +1700,7 @@ internal partial class Av1TileWriter Av1NeighborArrayUnit cr_dc_sign_level_coeff_na, Av1NeighborArrayUnit cb_dc_sign_level_coeff_na) { - EncodeTransformCoefficientsY( - pcs, - ec_ctx, - writer, - ref blk_ptr, - blockOrigin, - intraLumaDir, - planeBlockSize, - coefficientBuffer, - superblockIndex, - luma_dc_sign_level_coeff_na); - - EncodeTransformCoefficientsUv( + EncodeTransformCoefficientRegions( pcs, ec_ctx, writer, @@ -1722,6 +1710,7 @@ internal partial class Av1TileWriter planeBlockSize, coefficientBuffer, superblockIndex, + luma_dc_sign_level_coeff_na, cr_dc_sign_level_coeff_na, cb_dc_sign_level_coeff_na); } @@ -1751,14 +1740,6 @@ internal partial class Av1TileWriter int superblockIndex, Av1NeighborArrayUnit luma_dc_sign_level_coeff_na) { - ObuFrameHeader frameHeader = pcs.Parent.FrameHeader; - bool usesInterTransformSet = entropyCodingContext.MacroBlockModeInfo.Block.UseIntraBlockCopy; - Span lumaCoefficients = coefficientBuffer.GetPlaneSpan(superblockIndex, Av1Plane.Y); - Span lumaTransformBlocks = - coefficientBuffer.GetTransformBlockSpan(superblockIndex, Av1Plane.Y); - Av1TransformSize transformSize = entropyCodingContext.MacroBlockModeInfo.Block.TransformSize; - int transformBlockWidth = transformSize.Get4x4WideCount(); - int transformBlockHeight = transformSize.Get4x4HighCount(); Av1MacroBlockD macroBlock = entropyCodingContext.MacroBlock; int maximumBlocksWide = plane_bsize.GetWidth(); int maximumBlocksHigh = plane_bsize.GetHeight(); @@ -1777,66 +1758,33 @@ internal partial class Av1TileWriter int maximumUnitBlocksWide = Math.Min( Av1BlockSize.Block64x64.Get4x4WideCount(), maximumBlocksWide); + int maximumUnitBlocksHigh = Math.Min( Av1BlockSize.Block64x64.Get4x4HighCount(), maximumBlocksHigh); - // AV1 visits residuals in bounded 64x64 regions so transform order remains stable for 128x128 blocks. for (int regionRow = 0; regionRow < maximumBlocksHigh; regionRow += maximumUnitBlocksHigh) { - int unitHeight = Math.Min(maximumUnitBlocksHigh + regionRow, maximumBlocksHigh); + int unitBottom = Math.Min(regionRow + maximumUnitBlocksHigh, maximumBlocksHigh); for (int regionColumn = 0; regionColumn < maximumBlocksWide; regionColumn += maximumUnitBlocksWide) { - int unitWidth = Math.Min(maximumUnitBlocksWide + regionColumn, maximumBlocksWide); - for (int blockRow = regionRow; blockRow < unitHeight; blockRow += transformBlockHeight) - { - for (int blockColumn = regionColumn; blockColumn < unitWidth; blockColumn += transformBlockWidth) - { - int transformStateIndex = entropyCodingContext.CodedAreaSuperblock / - Av1EncoderCoefficientBuffer.TransformBlockUnitCoefficientCount; - ref Av1EncoderTransformBlockState transformBlock = ref lumaTransformBlocks[transformStateIndex]; - Point transformOrigin = blockOrigin + new Size( - blockColumn << Av1Constants.ModeInfoSizeLog2, - blockRow << Av1Constants.ModeInfoSizeLog2); - Span coefficients = lumaCoefficients[entropyCodingContext.CodedAreaSuperblock..]; - Av1TransformBlockContext blockContext = GetTransformBlockContexts( - Av1ComponentType.Luminance, - luma_dc_sign_level_coeff_na, - transformOrigin, - plane_bsize, - transformSize); - - Av1TransformType transformType = transformBlock.TransformType; - ushort endOfBlock = transformBlock.EndOfBlock; - if (endOfBlock == 0) - { - // Empty transform blocks use the canonical transform type even when mode decision retained another candidate. - transformType = transformBlock.TransformType = Av1TransformType.DctDct; - } - - int culLevelY = writer.WriteCoefficients( - transformSize, - transformType, - intraLumaDir, - coefficients, - Av1ComponentType.Luminance, - blockContext, - endOfBlock, - frameHeader.UseReducedTransformSet, - blk_ptr.FilterIntraMode, - usesInterTransformSet); - - int transformWidth = transformSize.GetWidth(); - int transformHeight = transformSize.GetHeight(); - luma_dc_sign_level_coeff_na.UnitModeWrite( - (byte)culLevelY, - transformOrigin, - new Size(transformWidth, transformHeight), - Av1NeighborArrayUnit.UnitMask.Top | Av1NeighborArrayUnit.UnitMask.Left); - - entropyCodingContext.CodedAreaSuperblock += transformWidth * transformHeight; - } - } + int unitRight = Math.Min(regionColumn + maximumUnitBlocksWide, maximumBlocksWide); + EncodeTransformCoefficientRegion( + pcs, + entropyCodingContext, + writer, + ref blk_ptr, + blockOrigin, + intraLumaDir, + plane_bsize, + Av1Plane.Y, + coefficientBuffer, + superblockIndex, + luma_dc_sign_level_coeff_na, + regionRow, + regionColumn, + unitBottom, + unitRight); } } } @@ -1874,26 +1822,10 @@ internal partial class Av1TileWriter return; } - ObuFrameHeader frameHeader = pcs.Parent.FrameHeader; - bool usesInterTransformSet = entropyCodingContext.MacroBlockModeInfo.Block.UseIntraBlockCopy; - Span blueCoefficients = coefficientBuffer.GetPlaneSpan(superblockIndex, Av1Plane.U); - Span redCoefficients = coefficientBuffer.GetPlaneSpan(superblockIndex, Av1Plane.V); - Span blueTransformBlocks = - coefficientBuffer.GetTransformBlockSpan(superblockIndex, Av1Plane.U); - Span redTransformBlocks = - coefficientBuffer.GetTransformBlockSpan(superblockIndex, Av1Plane.V); - int subsamplingX = colorConfig.SubSamplingX ? 1 : 0; int subsamplingY = colorConfig.SubSamplingY ? 1 : 0; Av1BlockSize chromaBlockSize = plane_bsize.GetSubsampled(colorConfig.SubSamplingX, colorConfig.SubSamplingY); Point chromaBlockOrigin = GetChromaBlockOrigin(blockOrigin, subsamplingX, subsamplingY); - Av1TransformSize chromaTransformSize = frameHeader.LosslessArray[entropyCodingContext.MacroBlockModeInfo.Block.SegmentId] - ? Av1TransformSize.Size4x4 - : plane_bsize.GetMaxUvTransformSize(colorConfig.SubSamplingX, colorConfig.SubSamplingY); - int transformBlockWidth = chromaTransformSize.Get4x4WideCount(); - int transformBlockHeight = chromaTransformSize.Get4x4HighCount(); - int transformWidth = chromaTransformSize.GetWidth(); - int transformHeight = chromaTransformSize.GetHeight(); Av1MacroBlockD macroBlock = entropyCodingContext.MacroBlock; int maximumBlocksWide = chromaBlockSize.GetWidth(); int maximumBlocksHigh = chromaBlockSize.GetHeight(); @@ -1911,81 +1843,272 @@ internal partial class Av1TileWriter maximumBlocksHigh >>= Av1Constants.ModeInfoSizeLog2; Av1BlockSize maximumUnitBlockSize = Av1BlockSize.Block64x64.GetSubsampled(colorConfig.SubSamplingX, colorConfig.SubSamplingY); + int maximumUnitBlocksWide = Math.Min(maximumUnitBlockSize.Get4x4WideCount(), maximumBlocksWide); int maximumUnitBlocksHigh = Math.Min(maximumUnitBlockSize.Get4x4HighCount(), maximumBlocksHigh); - int codedAreaStart = entropyCodingContext.CodedAreaSuperblockUv; - int codedAreaEnd = codedAreaStart; + for (int regionRow = 0; regionRow < maximumBlocksHigh; regionRow += maximumUnitBlocksHigh) + { + int unitBottom = Math.Min(regionRow + maximumUnitBlocksHigh, maximumBlocksHigh); + for (int regionColumn = 0; regionColumn < maximumBlocksWide; regionColumn += maximumUnitBlocksWide) + { + int unitRight = Math.Min(regionColumn + maximumUnitBlocksWide, maximumBlocksWide); + EncodeTransformCoefficientRegion( + pcs, + entropyCodingContext, + writer, + ref blk_ptr, + chromaBlockOrigin, + intraLumaDir, + plane_bsize, + Av1Plane.U, + coefficientBuffer, + superblockIndex, + cb_dc_sign_level_coeff_na, + regionRow, + regionColumn, + unitBottom, + unitRight); + + EncodeTransformCoefficientRegion( + pcs, + entropyCodingContext, + writer, + ref blk_ptr, + chromaBlockOrigin, + intraLumaDir, + plane_bsize, + Av1Plane.V, + coefficientBuffer, + superblockIndex, + cr_dc_sign_level_coeff_na, + regionRow, + regionColumn, + unitBottom, + unitRight); + } + } + } - // AV1 completes every transform in one chroma plane before advancing to the other plane. - // Both planes use the same coded-area positions because their transform geometry is identical. - for (int planeIndex = 0; planeIndex < 2; planeIndex++) + private static void EncodeTransformCoefficientRegions( + Av1PictureControlSet pcs, + Av1EntropyCodingContext entropyCodingContext, + Av1SymbolEncoder writer, + ref Av1EncoderBlockStruct block, + Point blockOrigin, + Av1PredictionMode intraLumaMode, + Av1BlockSize blockSize, + Av1EncoderCoefficientBuffer coefficientBuffer, + int superblockIndex, + Av1NeighborArrayUnit lumaCoefficientNeighbors, + Av1NeighborArrayUnit redCoefficientNeighbors, + Av1NeighborArrayUnit blueCoefficientNeighbors) + { + Av1MacroBlockD macroBlock = entropyCodingContext.MacroBlock; + int maximumBlocksWide = blockSize.GetWidth(); + int maximumBlocksHigh = blockSize.GetHeight(); + if (macroBlock.ToRightEdge < 0) + { + maximumBlocksWide += macroBlock.ToRightEdge >> 3; + } + + if (macroBlock.ToBottomEdge < 0) { - bool isBluePlane = planeIndex == 0; - Span planeCoefficients = isBluePlane ? blueCoefficients : redCoefficients; - Span planeTransformBlocks = - isBluePlane ? blueTransformBlocks : redTransformBlocks; + maximumBlocksHigh += macroBlock.ToBottomEdge >> 3; + } + + maximumBlocksWide >>= Av1Constants.ModeInfoSizeLog2; + maximumBlocksHigh >>= Av1Constants.ModeInfoSizeLog2; + int maximumUnitBlocksWide = Math.Min( + Av1BlockSize.Block64x64.Get4x4WideCount(), + maximumBlocksWide); + + int maximumUnitBlocksHigh = Math.Min( + Av1BlockSize.Block64x64.Get4x4HighCount(), + maximumBlocksHigh); - Av1NeighborArrayUnit coefficientNeighbors = - isBluePlane ? cb_dc_sign_level_coeff_na : cr_dc_sign_level_coeff_na; + ObuColorConfig colorConfig = pcs.Sequence.SequenceHeader.ColorConfig; + bool hasChroma = block.HasChroma && !colorConfig.IsMonochrome; + int subsamplingX = colorConfig.SubSamplingX ? 1 : 0; + int subsamplingY = colorConfig.SubSamplingY ? 1 : 0; + Point chromaBlockOrigin = GetChromaBlockOrigin(blockOrigin, subsamplingX, subsamplingY); - int codedArea = codedAreaStart; - for (int regionRow = 0; regionRow < maximumBlocksHigh; regionRow += maximumUnitBlocksHigh) + // Residual syntax is region-major, then plane-major. Keeping the three plane calls together + // prevents a 128x128 block from emitting later luma regions before earlier chroma regions. + for (int regionRow = 0; regionRow < maximumBlocksHigh; regionRow += maximumUnitBlocksHigh) + { + int unitBottom = Math.Min(regionRow + maximumUnitBlocksHigh, maximumBlocksHigh); + for (int regionColumn = 0; regionColumn < maximumBlocksWide; regionColumn += maximumUnitBlocksWide) { - int unitHeight = Math.Min(maximumUnitBlocksHigh + regionRow, maximumBlocksHigh); - for (int regionColumn = 0; regionColumn < maximumBlocksWide; regionColumn += maximumUnitBlocksWide) + int unitRight = Math.Min(regionColumn + maximumUnitBlocksWide, maximumBlocksWide); + EncodeTransformCoefficientRegion( + pcs, + entropyCodingContext, + writer, + ref block, + blockOrigin, + intraLumaMode, + blockSize, + Av1Plane.Y, + coefficientBuffer, + superblockIndex, + lumaCoefficientNeighbors, + regionRow, + regionColumn, + unitBottom, + unitRight); + + if (hasChroma) { - int unitWidth = Math.Min(maximumUnitBlocksWide + regionColumn, maximumBlocksWide); - for (int blockRow = regionRow; blockRow < unitHeight; blockRow += transformBlockHeight) - { - for (int blockColumn = regionColumn; blockColumn < unitWidth; blockColumn += transformBlockWidth) - { - int transformStateIndex = - codedArea / Av1EncoderCoefficientBuffer.TransformBlockUnitCoefficientCount; - - ref Av1EncoderTransformBlockState transformBlock = - ref planeTransformBlocks[transformStateIndex]; - - Point chromaOrigin = chromaBlockOrigin + new Size( - blockColumn << Av1Constants.ModeInfoSizeLog2, - blockRow << Av1Constants.ModeInfoSizeLog2); - - Span coefficients = planeCoefficients[codedArea..]; - Av1TransformBlockContext blockContext = GetTransformBlockContexts( - Av1ComponentType.Chroma, - coefficientNeighbors, - chromaOrigin, - chromaBlockSize, - chromaTransformSize); - - int culLevel = writer.WriteCoefficients( - chromaTransformSize, - transformBlock.TransformType, - intraLumaDir, - coefficients, - Av1ComponentType.Chroma, - blockContext, - transformBlock.EndOfBlock, - frameHeader.UseReducedTransformSet, - blk_ptr.FilterIntraMode, - usesInterTransformSet); - - coefficientNeighbors.UnitModeWrite( - (byte)culLevel, - chromaOrigin, - new Size(transformWidth, transformHeight), - Av1NeighborArrayUnit.UnitMask.Top | Av1NeighborArrayUnit.UnitMask.Left); - - codedArea += transformWidth * transformHeight; - } - } + int chromaRegionRow = regionRow >> subsamplingY; + int chromaRegionColumn = regionColumn >> subsamplingX; + int chromaUnitBottom = unitBottom >> subsamplingY; + int chromaUnitRight = unitRight >> subsamplingX; + EncodeTransformCoefficientRegion( + pcs, + entropyCodingContext, + writer, + ref block, + chromaBlockOrigin, + intraLumaMode, + blockSize, + Av1Plane.U, + coefficientBuffer, + superblockIndex, + blueCoefficientNeighbors, + chromaRegionRow, + chromaRegionColumn, + chromaUnitBottom, + chromaUnitRight); + + EncodeTransformCoefficientRegion( + pcs, + entropyCodingContext, + writer, + ref block, + chromaBlockOrigin, + intraLumaMode, + blockSize, + Av1Plane.V, + coefficientBuffer, + superblockIndex, + redCoefficientNeighbors, + chromaRegionRow, + chromaRegionColumn, + chromaUnitBottom, + chromaUnitRight); } } + } + } + + private static void EncodeTransformCoefficientRegion( + Av1PictureControlSet pcs, + Av1EntropyCodingContext entropyCodingContext, + Av1SymbolEncoder writer, + ref Av1EncoderBlockStruct block, + Point planeBlockOrigin, + Av1PredictionMode intraLumaMode, + Av1BlockSize lumaBlockSize, + Av1Plane plane, + Av1EncoderCoefficientBuffer coefficientBuffer, + int superblockIndex, + Av1NeighborArrayUnit coefficientNeighbors, + int regionRow, + int regionColumn, + int unitBottom, + int unitRight) + { + ObuFrameHeader frameHeader = pcs.Parent.FrameHeader; + ObuColorConfig colorConfig = pcs.Sequence.SequenceHeader.ColorConfig; + bool isLuma = plane == Av1Plane.Y; + Av1BlockSize planeBlockSize = isLuma + ? lumaBlockSize + : lumaBlockSize.GetSubsampled(colorConfig.SubSamplingX, colorConfig.SubSamplingY); - codedAreaEnd = codedArea; + Av1TransformSize transformSize = isLuma + ? entropyCodingContext.MacroBlockModeInfo.Block.TransformSize + : frameHeader.LosslessArray[entropyCodingContext.MacroBlockModeInfo.Block.SegmentId] + ? Av1TransformSize.Size4x4 + : lumaBlockSize.GetMaxUvTransformSize(colorConfig.SubSamplingX, colorConfig.SubSamplingY); + + int transformBlockWidth = transformSize.Get4x4WideCount(); + int transformBlockHeight = transformSize.Get4x4HighCount(); + int transformWidth = transformSize.GetWidth(); + int transformHeight = transformSize.GetHeight(); + bool usesInterTransformSet = entropyCodingContext.MacroBlockModeInfo.Block.UseIntraBlockCopy; + Av1ComponentType componentType = isLuma + ? Av1ComponentType.Luminance + : Av1ComponentType.Chroma; + + Span planeCoefficients = coefficientBuffer.GetPlaneSpan(superblockIndex, plane); + Span planeTransformBlocks = + coefficientBuffer.GetTransformBlockSpan(superblockIndex, plane); + + int codedArea = isLuma + ? entropyCodingContext.CodedAreaSuperblock + : entropyCodingContext.CodedAreaSuperblockUv; + + for (int blockRow = regionRow; blockRow < unitBottom; blockRow += transformBlockHeight) + { + for (int blockColumn = regionColumn; blockColumn < unitRight; blockColumn += transformBlockWidth) + { + int transformStateIndex = + codedArea / Av1EncoderCoefficientBuffer.TransformBlockUnitCoefficientCount; + + ref Av1EncoderTransformBlockState transformBlock = + ref planeTransformBlocks[transformStateIndex]; + + Point transformOrigin = planeBlockOrigin + new Size( + blockColumn << Av1Constants.ModeInfoSizeLog2, + blockRow << Av1Constants.ModeInfoSizeLog2); + + Span coefficients = planeCoefficients[codedArea..]; + Av1TransformBlockContext blockContext = GetTransformBlockContexts( + componentType, + coefficientNeighbors, + transformOrigin, + planeBlockSize, + transformSize); + + Av1TransformType transformType = transformBlock.TransformType; + if (isLuma && transformBlock.EndOfBlock == 0) + { + // Empty luma transforms carry no transform-type symbol, so retain the canonical state. + transformType = transformBlock.TransformType = Av1TransformType.DctDct; + } + + int culLevel = writer.WriteCoefficients( + transformSize, + transformType, + intraLumaMode, + coefficients, + componentType, + blockContext, + transformBlock.EndOfBlock, + frameHeader.UseReducedTransformSet, + block.FilterIntraMode, + usesInterTransformSet); + + coefficientNeighbors.UnitModeWrite( + (byte)culLevel, + transformOrigin, + new Size(transformWidth, transformHeight), + Av1NeighborArrayUnit.UnitMask.Top | Av1NeighborArrayUnit.UnitMask.Left); + + codedArea += transformWidth * transformHeight; + } } - entropyCodingContext.CodedAreaSuperblockUv = codedAreaEnd; + if (isLuma) + { + entropyCodingContext.CodedAreaSuperblock = codedArea; + } + else if (plane == Av1Plane.V) + { + // U and V share the same per-plane coded-area positions; advance only after V completes the region. + entropyCodingContext.CodedAreaSuperblockUv = codedArea; + } } /// diff --git a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs index d7d8a74125..c66b6641bc 100644 --- a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs +++ b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs @@ -357,6 +357,80 @@ public class Av1EncoderFrameTests Assert.Equal(new Size(Size, Size), decoded.Size); } + [Theory] + [InlineData(Yuv400)] + [InlineData(Yuv444)] + public void EncodeEffortTenSelectsOneHundredTwentyEightByOneHundredTwentyEightBlock(int colorFormatValue) + { + const int Size = 128; + Av1ColorFormat colorFormat = (Av1ColorFormat)colorFormatValue; + using Image source = new(Size, Size); + for (int y = 0; y < Size; y++) + { + Span row = source.Frames.RootFrame.PixelBuffer.DangerousGetRowSpan(y); + for (int x = 0; x < Size; x++) + { + row[x] = new Rgba32(180, 64, 220); + } + } + + using MemoryStream stream = new(); + ObuSequenceHeader sequenceHeader = Av1FrameEncoder.Encode( + Configuration.Default, + source.Frames.RootFrame, + stream, + CreateColorConfig(Av1BitDepth.EightBit, colorFormat), + qIndex: 4, + effort: 10); + + Assert.True(sequenceHeader.Use128x128Superblock); + byte[] payload = stream.ToArray(); + using Av1Decoder decoder = new(Configuration.Default); + using Image decoded = decoder.Decode(payload); + Av1FrameInfo frameInfo = Assert.IsType(decoder.FrameInfo); + for (int modeInfoY = 0; modeInfoY < 32; modeInfoY++) + { + for (int modeInfoX = 0; modeInfoX < 32; modeInfoX++) + { + Assert.Equal( + Av1BlockSize.Block128x128, + frameInfo.GetModeInfoAt(new Point(modeInfoX, modeInfoY)).BlockSize); + } + } + + Assert.Equal(new Size(Size, Size), decoded.Size); + } + + [Fact] + public void EncodeEffortTenSearchesHighBitDepthOneHundredTwentyEightRoot() + { + const int Size = 128; + using Image source = new(Size, Size); + for (int y = 0; y < Size; y++) + { + Span row = source.Frames.RootFrame.PixelBuffer.DangerousGetRowSpan(y); + for (int x = 0; x < Size; x++) + { + row[x] = new Rgba32(180, 64, 220); + } + } + + using MemoryStream stream = new(); + ObuSequenceHeader sequenceHeader = Av1FrameEncoder.Encode( + Configuration.Default, + source.Frames.RootFrame, + stream, + CreateColorConfig(Av1BitDepth.TwelveBit, Av1ColorFormat.Yuv444), + qIndex: 4, + effort: 10); + + Assert.True(sequenceHeader.Use128x128Superblock); + byte[] payload = stream.ToArray(); + using Av1Decoder decoder = new(Configuration.Default); + using Image decoded = decoder.Decode(payload); + Assert.Equal(new Size(Size, Size), decoded.Size); + } + [Theory] [InlineData(EightBit)] [InlineData(TenBit)] diff --git a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs index 0eee294fcb..62743aa865 100644 --- a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs +++ b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs @@ -618,19 +618,25 @@ public class Av1TransformBlockEncoderTests workspace.GetIntraBlockCopyWorkspace(); Assert.Equal( - (2 * Av1EncoderModeDecisionWorkspace.MaximumBlockDimension) + 1, + (2 * Av1Constants.MaxTransformSize) + 1, modeWorkspace.GetReferenceSamples(3).Length); Assert.Equal(Av1EncoderModeDecisionWorkspace.MaximumSampleCount, modeWorkspace.GetCandidateReconstruction(1).Length); Assert.Equal(Av1EncoderModeDecisionWorkspace.MaximumSampleCount, modeWorkspace.GetCandidateCoefficients(1).Length); + Assert.Equal(Av1EncoderModeDecisionWorkspace.MaximumTransformSampleCount, modeWorkspace.Prediction.Length); + Assert.Equal(Av1EncoderModeDecisionWorkspace.MaximumTransformSampleCount, modeWorkspace.Residual.Length); + Assert.Equal(Av1EncoderModeDecisionWorkspace.MaximumCandidateTransformBlockCount, modeWorkspace.CandidateTransformBlocks.Length); // CfL is unavailable above 32x32, so its scratch remains fixed while larger partitions are enabled. Assert.Equal(Av1ChromaFromLumaContext.BufferLength, modeWorkspace.ChromaFromLumaSamples.Length); Assert.Equal(Av1ChromaFromLumaMath.AlphaCandidateCount, modeWorkspace.GetChromaFromLumaRates(1).Length); Assert.Equal(Av1ChromaFromLumaMath.AlphaCandidateCount, modeWorkspace.GetChromaFromLumaDistortions(1).Length); - Assert.Equal(Av1EncoderModeDecisionWorkspace.MaximumSampleCount, paletteWorkspace.GetPrediction(1).Length); - Assert.Equal(Av1EncoderModeDecisionWorkspace.MaximumSampleCount, paletteWorkspace.AlternateIndices.Length); + int maximumPaletteSampleCount = + Av1BlockSize.Block64x64.GetWidth() * Av1BlockSize.Block64x64.GetHeight(); + + Assert.Equal(maximumPaletteSampleCount, paletteWorkspace.GetPrediction(1).Length); + Assert.Equal(maximumPaletteSampleCount, paletteWorkspace.AlternateIndices.Length); // Conventional mode search and IBC are sequential, so their typed views intentionally alias one owner region. modeWorkspace.GetReferenceSamples(0)[0] = 123;