diff --git a/HEIF_IMPLEMENTATION_PLAN.md b/HEIF_IMPLEMENTATION_PLAN.md index 721dccf951..9aeb7d6818 100644 --- a/HEIF_IMPLEMENTATION_PLAN.md +++ b/HEIF_IMPLEMENTATION_PLAN.md @@ -824,11 +824,11 @@ Encoder verification contract: - [~] A non-owning encoder-frame view now separates visible conversion regions from coded regions and performs complete left, top, right, bottom, and corner extension across each bordered plane. Current libaom uses 8-sample-aligned coded dimensions, a 32-sample-aligned luma stride with chroma stride derived from it, and a 64-pixel luma border for non-resized all-intra encoding. One operation-ready frame owner now rents the aligned Y, U, and V storage contiguously, exposes non-owning `Buffer2D` plane views, and returns the rent exactly once. A 4K 4:2:0 frame occupies about 13.0 MiB at 8-bit or 26.0 MiB at 10/12-bit; source and reconstruction therefore remain distinct frame owners rather than adding a full-frame copy. The corrected tests use this real ownership path and verify the exact 54 KiB 64x64 4:2:0 rent. The frame-encoder operation now instantiates matching source and reconstruction owners with ordinary `using` lifetimes and converts packed pixels directly into the source owner before extension. - [~] Temporal delimiter, sequence header, frame header, combined-frame tile-group writing, and an internal reduced-still-picture frame operation now exist locally. The remaining required metadata, padding, multi-tile, option, and public encoder paths are not complete. - [~] Implement superblock and partition analysis for every permitted block size and partition. The current baseline deliberately splits every in-frame node to 8x8 blocks and records decisions in current-libaom writer preorder; block-size selection and non-split partition analysis remain. -- [~] Implement intra mode search, chroma mode search, palette, filter intra, chroma-from-luma, and intra-block copy decisions. Live luma search now covers DC and the complete zero-angle H, V, Smooth, Paeth, Smooth-V, and Smooth-H family; diagonal modes, angle deltas, and the remaining decision families remain. +- [~] Implement intra mode search, chroma mode search, palette, filter intra, chroma-from-luma, and intra-block copy decisions. Live luma search now covers all 13 zero-angle base modes in current-libaom search order; angle deltas and the remaining decision families remain. - [ ] Implement inter mode search for bounded sequences, including reference selection and the decoder-supported inter tools. - [~] Current-libaom `av1_quantize_fp_no_qmatrix` arithmetic is implemented as a closed generic forward-quantizer family with Vector512, Vector256, Vector128, and scalar paths, raster-order output, coded 64-point coefficient limits, and scan-order EOB selection. Transform search, coefficient optimization, and lossless behavior remain. - [~] Implement real rate-distortion selection and make quality and effort change work, size, and output quality. The live zero-angle luma slice performs complete rate-distortion selection; quality mapping, effort-dependent pruning, and the remaining searches are not implemented. -- [~] Encoder rate accounting converts the entropy writer's live inverse cumulative distributions into current-libaom fixed-point symbol costs without allocating or duplicating probability state. Read-only luma-mode, directional-delta, filter-intra, chroma-mode, block-skip, transform-size, transform-block-skip, and complete transform-coefficient queries share the exact distributions mutated by the subsequent entropy write. Complete coefficient costing follows current libaom's optimized shape: it returns immediately for an empty transform, uses the EOB-specific base-range context, fuses magnitude, sign, base-range, and Golomb accounting into one reverse traversal, and combines repeated full base-range chunks instead of replaying each emitted symbol. Tile-lifetime level and context scratch is reused, the one-coefficient path neither clears nor initializes the forward-neighbor level map, and steady-state queries allocate nothing. Transform-size writing and costing share one subdivision-depth calculation, while shared closed symbol operations keep the writer and cost mappings for transform skip, transform type, and EOB syntax identical without forcing the estimator through the writer's slower two-pass coefficient traversal. The current-libaom fixed-point RD combiner preserves 64-bit distortion and rounds the weighted 1/512-bit rate at the required boundary. Its key-frame multiplier follows libaom's squared DC-quantizer formula and exact 10/12-bit normalization. Live final-block selection now evaluates DC, H, V, Smooth, Paeth, Smooth-V, and Smooth-H in current libaom's base-mode search order with complete transform, quantization, reconstruction, coefficient rate, and normalized pixel-domain distortion. Missing top or left edges and the Paeth corner are synthesized from the closest coded edge or the bit-depth midpoint offsets using the same rules as current libaom, so tile-edge modes are evaluated instead of incorrectly removed. The tile writer invokes this stack-only selector after mapping current neighbors and immediately before writing each block, so later decisions see reconstructed samples, coefficient contexts, and CDF updates from every preceding block. Block skip is read only after the callback has combined every coded plane. Candidate scratch remains one 8x8 reconstruction and one 8x8 coefficient span on the stack; only a newly winning candidate is copied into retained frame storage. Six independent production fixtures force each added predictor from already reconstructed top, left, and corner regions. The exact net11 Release rebuild remains at 1,005 warnings and zero errors, 36 focused predictor, block, mode-decision, and frame-operation cases pass, all 8,905 HEIF/AV1 namespace cases pass, and current-main `aomdec` accepts all 29 emitted 8/10/12-bit 4:0:0, 4:2:0, 4:2:2, and 4:4:4 constant or gradient payloads. Remaining mode decision work includes the six diagonal base modes and their angle deltas, transform-size/type search, chroma-mode search, partition search, full block-skip RD comparison, and effort-dependent pruning. +- [~] Encoder rate accounting converts the entropy writer's live inverse cumulative distributions into current-libaom fixed-point symbol costs without allocating or duplicating probability state. Read-only luma-mode, directional-delta, filter-intra, chroma-mode, block-skip, transform-size, transform-block-skip, and complete transform-coefficient queries share the exact distributions mutated by the subsequent entropy write. Complete coefficient costing follows current libaom's optimized shape: it returns immediately for an empty transform, uses the EOB-specific base-range context, fuses magnitude, sign, base-range, and Golomb accounting into one reverse traversal, and combines repeated full base-range chunks instead of replaying each emitted symbol. Tile-lifetime level and context scratch is reused, the one-coefficient path neither clears nor initializes the forward-neighbor level map, and steady-state queries allocate nothing. Transform-size writing and costing share one subdivision-depth calculation, while shared closed symbol operations keep the writer and cost mappings for transform skip, transform type, and EOB syntax identical without forcing the estimator through the writer's slower two-pass coefficient traversal. The current-libaom fixed-point RD combiner preserves 64-bit distortion and rounds the weighted 1/512-bit rate at the required boundary. Its key-frame multiplier follows libaom's squared DC-quantizer formula and exact 10/12-bit normalization. Live final-block selection now evaluates all 13 zero-angle base modes in current libaom's search order with complete transform, quantization, reconstruction, coefficient rate, and normalized pixel-domain distortion. Both prepared reference edges retain the common-corner prefix and 16 samples required by 8x8 directional prediction. A shared encoder/decoder availability calculation selects reconstructed top-right and bottom-left extensions according to tile, frame, superblock, and block reconstruction order; unavailable extensions repeat the nearest coded endpoint. Missing top or left edges retain current libaom's perpendicular-sample and bit-depth-midpoint rules. Directional prediction reuses transform workspace for zone-three transposition before the transform overwrites it, keeping candidate evaluation allocation-free. The tile writer invokes this stack-only selector after mapping current neighbors and immediately before writing each block, so later decisions see reconstructed samples, coefficient contexts, and CDF updates from every preceding block. Block skip is read only after the callback has combined every coded plane. Candidate scratch remains one 8x8 reconstruction and one 8x8 coefficient span on the stack; only a newly winning candidate is copied into retained frame storage. Independent production fixtures force every added predictor and prove that diagonal modes consume available top-right and bottom-left extensions. The exact net11 Release rebuild remains at 1,005 warnings and zero errors, 43 focused predictor, block, mode-decision, and decoder cases pass, all 8,912 HEIF/AV1 namespace cases pass, and current-main `aomdec` accepts all 29 emitted 8/10/12-bit 4:0:0, 4:2:0, 4:2:2, and 4:4:4 constant or gradient payloads. Remaining mode decision work includes directional angle deltas, transform-size/type search, chroma-mode search, partition search, full block-skip RD comparison, and effort-dependent pruning. - [~] The tile writer now publishes one packed coefficient context per covered 4x4 edge unit and derives luma/chroma skip plus DC-sign contexts from the complete transform edges using current-libaom units. Partition, transform, and coefficient neighbor state retains only the above and left context regions used by current libaom; the unused third top-left region, its granularity state, and its unused sentinel are removed. One picture owner now packs segmentation plus every tile's partition, luma, chroma, and transform edges into one clean byte allocation with typed non-owning views; together with the separately typed packed mode-information owner, the complete picture state uses two allocator rents rather than seven. Exact aligned lengths, clean initialization, and balanced exactly-once returns are covered in Release. Multi-tile payload ownership and verified CDF update behavior remain. - [~] Encoder mode information now uses a frame-owned integer alias grid over a packed 8-byte value allocation, matching current libaom's `mi_grid_base` and `mi_alloc` relationship without a managed object or reference per 4x4 entry. The visible dimensions are aligned to eight luma samples, the grid stride and allocated row count are aligned to 32 mode-information units, and optional 8x8 allocation granularity reduces the value store in both dimensions exactly as current libaom does. One clean ImageSharp byte owner contains both independently typed regions, reducing libaom's two allocation lifetimes to one without a copy. At 4K, the 4x4 layout occupies about 6.0 MiB in total; the 8x8 layout occupies about 3.0 MiB. Exact geometry, clean allocation, typed lengths, aligned mapping, untouched row padding, and exactly-once return pass 4 of 4 direct net11 VSTest cases in Release. Every coded 4x4 cell covered by square, rectangular, or clipped edge blocks maps to its owning allocation entry before context-dependent symbols are written. Packed syntax, relative neighbor lookup, full block mapping, writer traversal, entropy, and OBU coverage pass 1,947 of 1,947 direct net11 VSTest cases in Release; complete mode decision still remains. - [~] The final-block decision workspace uses one reusable 10.3 KiB ImageSharp allocator owner. It contains 1,024 explicitly packed 10-byte final-block entries and the 341 preorder partition bytes required by a complete 128x128-through-8x8 quadtree, replacing separate managed arrays. Construction and the explicit per-superblock reset initialize every syntax field, including the nonzero sentinel that disables filter-intra prediction; pooled palette, quantizer, prediction, and partition bytes cannot leak into the next decision pass. Exact allocation, size, initialization, reset, return, repeated-run, writer, entropy, and OBU coverage pass 1,957 of 1,957 direct net11 VSTest cases in Release; complete mode decision still remains. diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs index d9b44f19a1..17119de874 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs @@ -26,7 +26,13 @@ internal static partial class Av1IntraSuperblockEncoder Av1PredictionMode.Smooth, Av1PredictionMode.Paeth, Av1PredictionMode.SmoothVertical, - Av1PredictionMode.SmoothHorizontal + Av1PredictionMode.SmoothHorizontal, + Av1PredictionMode.Directional135Degrees, + Av1PredictionMode.Directional203Degrees, + Av1PredictionMode.Directional157Degrees, + Av1PredictionMode.Directional67Degrees, + Av1PredictionMode.Directional113Degrees, + Av1PredictionMode.Directional45Degrees ]; /// @@ -254,18 +260,51 @@ internal static partial class Av1IntraSuperblockEncoder Buffer2DRegion reconstructionPlane = this.reconstruction.GetPlane(Av1Plane.Y); bool hasLeft = macroBlock.IsLeftAvailable; bool hasAbove = macroBlock.IsUpAvailable; - Span aboveStorage = stackalloc TSample[9]; + int modeInfoRow = blockOrigin.Y >> Av1Constants.ModeInfoSizeLog2; + int modeInfoColumn = blockOrigin.X >> Av1Constants.ModeInfoSizeLog2; + bool rightAvailable = modeInfoColumn + TransformSize.Get4x4WideCount() < macroBlock.Tile.ModeInfoColumnEnd; + bool bottomAvailable = modeInfoRow + TransformSize.Get4x4HighCount() < macroBlock.Tile.ModeInfoRowEnd; + bool hasTopRight = Av1IntraReferenceAvailability.HasTopRight( + this.picture.Sequence.SequenceHeader.SuperblockSize, + BlockSize, + modeInfoRow, + modeInfoColumn, + hasAbove, + rightAvailable, + Av1PartitionType.None, + TransformSize, + 0, + 0, + 0, + 0); + + bool hasBottomLeft = Av1IntraReferenceAvailability.HasBottomLeft( + this.picture.Sequence.SequenceHeader.SuperblockSize, + BlockSize, + modeInfoRow, + modeInfoColumn, + bottomAvailable, + hasLeft, + Av1PartitionType.None, + TransformSize, + 0, + 0, + 0, + 0); + + Span aboveStorage = stackalloc TSample[17]; Span above = aboveStorage[1..]; - Span left = stackalloc TSample[8]; + Span leftStorage = stackalloc TSample[17]; + Span left = leftStorage[1..]; if (hasAbove) { - reconstructionPlane.DangerousGetRowSpan(blockOrigin.Y - 1).Slice(blockOrigin.X, 8).CopyTo(above); + reconstructionPlane.DangerousGetRowSpan(blockOrigin.Y - 1).Slice(blockOrigin.X, 8).CopyTo(above[..8]); } if (hasLeft) { - for (int row = 0; row < left.Length; row++) + for (int row = 0; row < 8; row++) { left[row] = reconstructionPlane.DangerousGetRowSpan(blockOrigin.Y + row)[blockOrigin.X - 1]; } @@ -277,17 +316,38 @@ internal static partial class Av1IntraSuperblockEncoder // available uses the asymmetric midpoint offsets that distinguish top from left. if (!hasAbove) { - above.Fill(hasLeft ? left[0] : TOperator.CreateSample(midpoint - 1)); + above[..8].Fill(hasLeft ? left[0] : TOperator.CreateSample(midpoint - 1)); } if (!hasLeft) { - left.Fill(hasAbove ? above[0] : TOperator.CreateSample(midpoint + 1)); + left[..8].Fill(hasAbove ? above[0] : TOperator.CreateSample(midpoint + 1)); + } + + if (hasTopRight) + { + reconstructionPlane.DangerousGetRowSpan(blockOrigin.Y - 1).Slice(blockOrigin.X + 8, 8).CopyTo(above[8..]); + } + else + { + above[8..].Fill(above[7]); } - // Paeth addresses the common corner immediately before the prepared top edge. When an edge is - // unavailable AV1 derives that corner from the closest coded edge, preserving tile independence. - aboveStorage[0] = hasAbove && hasLeft + if (hasBottomLeft) + { + for (int row = 8; row < 16; row++) + { + left[row] = reconstructionPlane.DangerousGetRowSpan(blockOrigin.Y + row)[blockOrigin.X - 1]; + } + } + else + { + left[8..].Fill(left[7]); + } + + // Zone-two projection and Paeth address the common corner immediately before both prepared edges. + // When an edge is unavailable AV1 derives that corner from the closest coded edge. + TSample corner = hasAbove && hasLeft ? reconstructionPlane.DangerousGetRowSpan(blockOrigin.Y - 1)[blockOrigin.X - 1] : hasAbove ? above[0] @@ -295,6 +355,9 @@ internal static partial class Av1IntraSuperblockEncoder ? left[0] : TOperator.CreateSample(midpoint); + aboveStorage[0] = corner; + leftStorage[0] = corner; + Av1TransformBlockContext blockContext = Av1TileWriter.GetTransformBlockContexts( Av1ComponentType.Luminance, this.picture.LuminanceDcSignLevelCoefficientNeighbors[tileIndex], diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1TransformBlockEncoder.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1TransformBlockEncoder.cs index 819fadaad1..a7e53c9228 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1TransformBlockEncoder.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1TransformBlockEncoder.cs @@ -352,6 +352,23 @@ internal static class Av1TransformBlockEncoder { Av1DcIntraPredictor.Predict(hasLeft, hasAbove, reconstruction, reconstructionStride, above, left, width, height); } + else if (mode.IsDirectional()) + { + // The current encoder disables intra-edge filtering in sequence syntax. Reusing transform workspace for + // zone-three transposition keeps directional prediction allocation-free before the transform overwrites it. + Span directionalScratch = MemoryMarshal.AsBytes(workspace.TransformWorkspace)[..(width * height)]; + + Av1DirectionalIntraPredictor.Predict( + reconstruction, + reconstructionStride, + transformSize, + above, + left, + false, + false, + mode.ToAngle(), + directionalScratch); + } else { Av1NonDirectionalIntraPredictorBase.GetPredictor(mode) @@ -450,6 +467,21 @@ internal static class Av1TransformBlockEncoder height, bitDepth.GetBitCount()); } + else if (mode.IsDirectional()) + { + Span directionalScratch = MemoryMarshal.Cast(workspace.TransformWorkspace)[..(width * height)]; + + Av1DirectionalIntraPredictor.Predict( + signedReconstruction, + reconstructionStride, + transformSize, + signedAbove, + signedLeft, + false, + false, + mode.ToAngle(), + directionalScratch); + } else { Av1NonDirectionalIntraPredictorBase.GetPredictor(mode) diff --git a/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1IntraReferenceAvailability.cs b/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1IntraReferenceAvailability.cs new file mode 100644 index 0000000000..4c839b8381 --- /dev/null +++ b/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1IntraReferenceAvailability.cs @@ -0,0 +1,191 @@ +// Copyright (c) Six Labors. +// Licensed under the Six Labors Split License. + +using System; +using SixLabors.ImageSharp.Formats.Heif.Av1.Transform; + +namespace SixLabors.ImageSharp.Formats.Heif.Av1.Prediction; + +/// +/// Determines whether extended intra-prediction references have already been reconstructed. +/// +internal static class Av1IntraReferenceAvailability +{ + /// + /// Determines whether every bottom-left reference sample required by a transform is already reconstructed. + /// + /// The sequence superblock size. + /// The containing block size in the current plane's geometry. + /// The containing block row in 4-by-4 mode-information units. + /// The containing block column in 4-by-4 mode-information units. + /// A value indicating whether the required rows remain inside the frame and tile. + /// A value indicating whether reconstructed samples exist immediately to the left. + /// The partition type that determines reconstruction order. + /// The transform size whose extended edge is required. + /// The transform row offset within the containing block. + /// The transform column offset within the containing block. + /// The horizontal chroma subsampling shift. + /// The vertical chroma subsampling shift. + /// when the bottom-left reference extension is available; otherwise, . + public static bool HasBottomLeft(Av1BlockSize superblockSize, Av1BlockSize blockSize, int modeInfoRow, int modeInfoColumn, bool bottomAvailable, bool haveLeft, Av1PartitionType partition, Av1TransformSize transformSize, int blockModeInfoRowOffset, int blockModeInfoColumnOffset, int subX, int subY) + { + if (!bottomAvailable || !haveLeft) + { + return false; + } + + // A 128-wide block is reconstructed as two 64-wide regions in raster order, + // so the right half can consume references that already belong to the left half. + if (blockSize.GetWidth() > 64 && blockModeInfoColumnOffset > 0) + { + int block64WidthInUnits = Av1BlockSize.Block64x64.Get4x4WideCount(); + int planeBlockWidthInUnits64 = block64WidthInUnits >> subX; + int columnOffset64 = blockModeInfoColumnOffset % planeBlockWidthInUnits64; + if (columnOffset64 == 0) + { + // We are at the left edge of top-right or bottom-right 64x* block. + int block64HeightInUnits = Av1BlockSize.Block64x64.Get4x4HighCount(); + int planeBlockHeightInUnits64 = block64HeightInUnits >> subY; + int rowOffset64 = blockModeInfoRowOffset % planeBlockHeightInUnits64; + int planeBlockHeightInUnits = Math.Min(blockSize.Get4x4HighCount() >> subY, planeBlockHeightInUnits64); + + // Check if all bottom-left pixels are in the left 64x* block (which is + // already coded). + return rowOffset64 + transformSize.Get4x4HighCount() < planeBlockHeightInUnits; + } + } + + if (blockModeInfoColumnOffset > 0) + { + // Bottom-left pixels are in the bottom-left block, which is not available. + return false; + } + else + { + int blockHeightInUnits = blockSize.GetHeight() >> Av1TransformSize.Size4x4.GetBlockHeightLog2(); + int planeBlockHeightInUnits = Math.Max(blockHeightInUnits >> subY, 1); + int bottomLeftUnitCount = transformSize.Get4x4HighCount(); + + // All bottom-left pixels are in the left block, which is already available. + if (blockModeInfoRowOffset + bottomLeftUnitCount < planeBlockHeightInUnits) + { + return true; + } + + int blockWidthInModeInfoLog2 = blockSize.Get4x4WidthLog2(); + int blockHeightInModeInfoLog2 = blockSize.Get4x4HeightLog2(); + int superblockModeInfoSize = superblockSize.Get4x4HighCount(); + int blockRowInSuperblock = (modeInfoRow & (superblockModeInfoSize - 1)) >> blockHeightInModeInfoLog2; + int blockColumnInSuperblock = (modeInfoColumn & (superblockModeInfoSize - 1)) >> blockWidthInModeInfoLog2; + + // Leftmost column of superblock: so bottom-left pixels maybe in the left + // and/or bottom-left superblocks. But only the left superblock is + // available, so check if all required pixels fall in that superblock. + if (blockColumnInSuperblock == 0) + { + int blockStartRowOffset = blockRowInSuperblock << (blockHeightInModeInfoLog2 + Av1Constants.ModeInfoSizeLog2 - Av1TransformSize.Size4x4.GetBlockWidthLog2()) >> subY; + int rowOffsetInSuperblock = blockStartRowOffset + blockModeInfoRowOffset; + int superblockHeightInUnits = superblockModeInfoSize >> subY; + return rowOffsetInSuperblock + bottomLeftUnitCount < superblockHeightInUnits; + } + + // Bottom row of superblock (and not the leftmost column): so bottom-left + // pixels fall in the bottom superblock, which is not available yet. + if (((blockRowInSuperblock + 1) << blockHeightInModeInfoLog2) >= superblockModeInfoSize) + { + return false; + } + + // General case (neither leftmost column nor bottom row): check if the + // bottom-left block is coded before the current block. + int thisBlockIndex = ((blockRowInSuperblock + 0) << (Av1Constants.MaxSuperBlockSizeLog2 - Av1Constants.ModeInfoSizeLog2 - blockWidthInModeInfoLog2)) + blockColumnInSuperblock + 0; + return Av1BottomRightTopLeftConstants.HasBottomLeft(partition, blockSize, thisBlockIndex); + } + } + + /// + /// Determines whether every top-right reference sample required by a transform is already reconstructed. + /// + /// The sequence superblock size. + /// The containing block size in the current plane's geometry. + /// The containing block row in 4-by-4 mode-information units. + /// The containing block column in 4-by-4 mode-information units. + /// A value indicating whether reconstructed samples exist immediately above. + /// A value indicating whether the required columns remain inside the frame and tile. + /// The partition type that determines reconstruction order. + /// The transform size whose extended edge is required. + /// The transform row offset within the containing block. + /// The transform column offset within the containing block. + /// The horizontal chroma subsampling shift. + /// The vertical chroma subsampling shift. + /// when the top-right reference extension is available; otherwise, . + public static bool HasTopRight(Av1BlockSize superblockSize, Av1BlockSize blockSize, int modeInfoRow, int modeInfoColumn, bool haveTop, bool rightAvailable, Av1PartitionType partition, Av1TransformSize transformSize, int blockModeInfoRowOffset, int blockModeInfoColumnOffset, int subX, int subY) + { + if (!haveTop || !rightAvailable) + { + return false; + } + + int blockWideInUnits = blockSize.GetWidth() >> 2; + int planeBlockWidthInUnits = Math.Max(blockWideInUnits >> subX, 1); + int topRightUnitCount = transformSize.Get4x4WideCount(); + + if (blockModeInfoRowOffset > 0) + { + // Transforms below the first row obtain their top edge from the containing block, + // so only the reconstructed width to their right constrains availability. + if (blockSize.GetWidth() > 64) + { + // Special case: For 128x128 blocks, the transform unit whose + // top-right corner is at the center of the block does in fact have + // pixels available at its top-right corner. + int block64WidthInUnits = Av1BlockSize.Block64x64.Get4x4WideCount(); + int block64HeightInUnits = Av1BlockSize.Block64x64.Get4x4HighCount(); + if (blockModeInfoRowOffset == block64HeightInUnits >> subY && + blockModeInfoColumnOffset + topRightUnitCount == block64WidthInUnits >> subX) + { + return true; + } + + int planeBlockWidthInUnits64 = block64WidthInUnits >> subX; + int blockModeInfoColumnOffset64 = blockModeInfoColumnOffset % planeBlockWidthInUnits64; + return blockModeInfoColumnOffset64 + topRightUnitCount < planeBlockWidthInUnits64; + } + + return blockModeInfoColumnOffset + topRightUnitCount < planeBlockWidthInUnits; + } + else + { + // All top-right pixels are in the block above, which is already available. + if (blockModeInfoColumnOffset + topRightUnitCount < planeBlockWidthInUnits) + { + return true; + } + + int blockWidthInModeInfoLog2 = blockSize.Get4x4WidthLog2(); + int blockHeightInModeInfeLog2 = blockSize.Get4x4HeightLog2(); + int superBlockModeInfoSize = superblockSize.Get4x4HighCount(); + int blockRowInSuperblock = (modeInfoRow & (superBlockModeInfoSize - 1)) >> blockHeightInModeInfeLog2; + int blockColumnInSuperBlock = (modeInfoColumn & (superBlockModeInfoSize - 1)) >> blockWidthInModeInfoLog2; + + // Top row of superblock: so top-right pixels are in the top and/or + // top-right superblocks, both of which are already available. + if (blockRowInSuperblock == 0) + { + return true; + } + + // Rightmost column of superblock (and not the top row): so top-right pixels + // fall in the right superblock, which is not available yet. + if (((blockColumnInSuperBlock + 1) << blockWidthInModeInfoLog2) >= superBlockModeInfoSize) + { + return false; + } + + // General case (neither top row nor rightmost column): check if the + // top-right block is coded before the current block. + int thisBlockIndex = ((blockRowInSuperblock + 0) << (Av1Constants.MaxSuperBlockSizeLog2 - Av1Constants.ModeInfoSizeLog2 - blockWidthInModeInfoLog2)) + blockColumnInSuperBlock + 0; + return Av1BottomRightTopLeftConstants.HasTopRight(partition, blockSize, thisBlockIndex); + } + } +} diff --git a/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1PredictionDecoder.cs b/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1PredictionDecoder.cs index 75ce908391..a4319d945b 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1PredictionDecoder.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1PredictionDecoder.cs @@ -518,7 +518,7 @@ internal sealed class Av1PredictionDecoder // Chroma prediction geometry cannot be smaller than 4 by 4 after subsampling. blockSize = ScaleChromaBlockSize(blockSize, subX == 1, subY == 1); - bool haveTopRight = IntraHasTopRight( + bool haveTopRight = Av1IntraReferenceAvailability.HasTopRight( this.sequenceHeader.SuperblockSize, blockSize, modeInfoRow, @@ -531,7 +531,7 @@ internal sealed class Av1PredictionDecoder blockModeInfoColumnOffset, subX, subY); - bool haveBottomLeft = IntraHasBottomLeft( + bool haveBottomLeft = Av1IntraReferenceAvailability.HasBottomLeft( this.sequenceHeader.SuperblockSize, blockSize, modeInfoRow, @@ -662,184 +662,6 @@ internal sealed class Av1PredictionDecoder return bs; } - /// - /// Determines whether every bottom-left reference sample required by a transform is already reconstructed. - /// - /// The sequence superblock size. - /// The containing block size in the current plane's geometry. - /// The containing block row in 4-by-4 mode-information units. - /// The containing block column in 4-by-4 mode-information units. - /// A value indicating whether the required rows remain inside the frame and tile. - /// A value indicating whether reconstructed samples exist immediately to the left. - /// The partition type that determines reconstruction order. - /// The transform size whose extended edge is required. - /// The transform row offset within the containing block. - /// The transform column offset within the containing block. - /// The horizontal chroma subsampling shift. - /// The vertical chroma subsampling shift. - /// when the bottom-left reference extension is available; otherwise, . - private static bool IntraHasBottomLeft(Av1BlockSize superblockSize, Av1BlockSize blockSize, int modeInfoRow, int modeInfoColumn, bool bottomAvailable, bool haveLeft, Av1PartitionType partition, Av1TransformSize transformSize, int blockModeInfoRowOffset, int blockModeInfoColumnOffset, int subX, int subY) - { - if (!bottomAvailable || !haveLeft) - { - return false; - } - - // A 128-wide block is reconstructed as two 64-wide regions in raster order, - // so the right half can consume references that already belong to the left half. - if (blockSize.GetWidth() > 64 && blockModeInfoColumnOffset > 0) - { - int block64WidthInUnits = Av1BlockSize.Block64x64.Get4x4WideCount(); - int planeBlockWidthInUnits64 = block64WidthInUnits >> subX; - int columnOffset64 = blockModeInfoColumnOffset % planeBlockWidthInUnits64; - if (columnOffset64 == 0) - { - // We are at the left edge of top-right or bottom-right 64x* block. - int block64HeightInUnits = Av1BlockSize.Block64x64.Get4x4HighCount(); - int planeBlockHeightInUnits64 = block64HeightInUnits >> subY; - int rowOffset64 = blockModeInfoRowOffset % planeBlockHeightInUnits64; - int planeBlockHeightInUnits = Math.Min(blockSize.Get4x4HighCount() >> subY, planeBlockHeightInUnits64); - - // Check if all bottom-left pixels are in the left 64x* block (which is - // already coded). - return rowOffset64 + transformSize.Get4x4HighCount() < planeBlockHeightInUnits; - } - } - - if (blockModeInfoColumnOffset > 0) - { - // Bottom-left pixels are in the bottom-left block, which is not available. - return false; - } - else - { - int blockHeightInUnits = blockSize.GetHeight() >> Av1TransformSize.Size4x4.GetBlockHeightLog2(); - int planeBlockHeightInUnits = Math.Max(blockHeightInUnits >> subY, 1); - int bottomLeftUnitCount = transformSize.Get4x4HighCount(); - - // All bottom-left pixels are in the left block, which is already available. - if (blockModeInfoRowOffset + bottomLeftUnitCount < planeBlockHeightInUnits) - { - return true; - } - - int blockWidthInModeInfoLog2 = blockSize.Get4x4WidthLog2(); - int blockHeightInModeInfoLog2 = blockSize.Get4x4HeightLog2(); - int superblockModeInfoSize = superblockSize.Get4x4HighCount(); - int blockRowInSuperblock = (modeInfoRow & (superblockModeInfoSize - 1)) >> blockHeightInModeInfoLog2; - int blockColumnInSuperblock = (modeInfoColumn & (superblockModeInfoSize - 1)) >> blockWidthInModeInfoLog2; - - // Leftmost column of superblock: so bottom-left pixels maybe in the left - // and/or bottom-left superblocks. But only the left superblock is - // available, so check if all required pixels fall in that superblock. - if (blockColumnInSuperblock == 0) - { - int blockStartRowOffset = blockRowInSuperblock << (blockHeightInModeInfoLog2 + Av1Constants.ModeInfoSizeLog2 - Av1TransformSize.Size4x4.GetBlockWidthLog2()) >> subY; - int rowOffsetInSuperblock = blockStartRowOffset + blockModeInfoRowOffset; - int superblockHeightInUnits = superblockModeInfoSize >> subY; - return rowOffsetInSuperblock + bottomLeftUnitCount < superblockHeightInUnits; - } - - // Bottom row of superblock (and not the leftmost column): so bottom-left - // pixels fall in the bottom superblock, which is not available yet. - if (((blockRowInSuperblock + 1) << blockHeightInModeInfoLog2) >= superblockModeInfoSize) - { - return false; - } - - // General case (neither leftmost column nor bottom row): check if the - // bottom-left block is coded before the current block. - int thisBlockIndex = ((blockRowInSuperblock + 0) << (Av1Constants.MaxSuperBlockSizeLog2 - Av1Constants.ModeInfoSizeLog2 - blockWidthInModeInfoLog2)) + blockColumnInSuperblock + 0; - return Av1BottomRightTopLeftConstants.HasBottomLeft(partition, blockSize, thisBlockIndex); - } - } - - /// - /// Determines whether every top-right reference sample required by a transform is already reconstructed. - /// - /// The sequence superblock size. - /// The containing block size in the current plane's geometry. - /// The containing block row in 4-by-4 mode-information units. - /// The containing block column in 4-by-4 mode-information units. - /// A value indicating whether reconstructed samples exist immediately above. - /// A value indicating whether the required columns remain inside the frame and tile. - /// The partition type that determines reconstruction order. - /// The transform size whose extended edge is required. - /// The transform row offset within the containing block. - /// The transform column offset within the containing block. - /// The horizontal chroma subsampling shift. - /// The vertical chroma subsampling shift. - /// when the top-right reference extension is available; otherwise, . - private static bool IntraHasTopRight(Av1BlockSize superblockSize, Av1BlockSize blockSize, int modeInfoRow, int modeInfoColumn, bool haveTop, bool rightAvailable, Av1PartitionType partition, Av1TransformSize transformSize, int blockModeInfoRowOffset, int blockModeInfoColumnOffset, int subX, int subY) - { - if (!haveTop || !rightAvailable) - { - return false; - } - - int blockWideInUnits = blockSize.GetWidth() >> 2; - int planeBlockWidthInUnits = Math.Max(blockWideInUnits >> subX, 1); - int topRightUnitCount = transformSize.Get4x4WideCount(); - - if (blockModeInfoRowOffset > 0) - { - // Transforms below the first row obtain their top edge from the containing block, - // so only the reconstructed width to their right constrains availability. - if (blockSize.GetWidth() > 64) - { - // Special case: For 128x128 blocks, the transform unit whose - // top-right corner is at the center of the block does in fact have - // pixels available at its top-right corner. - int block64WidthInUnits = Av1BlockSize.Block64x64.Get4x4WideCount(); - int block64HeightInUnits = Av1BlockSize.Block64x64.Get4x4HighCount(); - if (blockModeInfoRowOffset == block64HeightInUnits >> subY && - blockModeInfoColumnOffset + topRightUnitCount == block64WidthInUnits >> subX) - { - return true; - } - - int planeBlockWidthInUnits64 = block64WidthInUnits >> subX; - int blockModeInfoColumnOffset64 = blockModeInfoColumnOffset % planeBlockWidthInUnits64; - return blockModeInfoColumnOffset64 + topRightUnitCount < planeBlockWidthInUnits64; - } - - return blockModeInfoColumnOffset + topRightUnitCount < planeBlockWidthInUnits; - } - else - { - // All top-right pixels are in the block above, which is already available. - if (blockModeInfoColumnOffset + topRightUnitCount < planeBlockWidthInUnits) - { - return true; - } - - int blockWidthInModeInfoLog2 = blockSize.Get4x4WidthLog2(); - int blockHeightInModeInfeLog2 = blockSize.Get4x4HeightLog2(); - int superBlockModeInfoSize = superblockSize.Get4x4HighCount(); - int blockRowInSuperblock = (modeInfoRow & (superBlockModeInfoSize - 1)) >> blockHeightInModeInfeLog2; - int blockColumnInSuperBlock = (modeInfoColumn & (superBlockModeInfoSize - 1)) >> blockWidthInModeInfoLog2; - - // Top row of superblock: so top-right pixels are in the top and/or - // top-right superblocks, both of which are already available. - if (blockRowInSuperblock == 0) - { - return true; - } - - // Rightmost column of superblock (and not the top row): so top-right pixels - // fall in the right superblock, which is not available yet. - if (((blockColumnInSuperBlock + 1) << blockWidthInModeInfoLog2) >= superBlockModeInfoSize) - { - return false; - } - - // General case (neither top row nor rightmost column): check if the - // top-right block is coded before the current block. - int thisBlockIndex = ((blockRowInSuperblock + 0) << (Av1Constants.MaxSuperBlockSizeLog2 - Av1Constants.ModeInfoSizeLog2 - blockWidthInModeInfoLog2)) + blockColumnInSuperBlock + 0; - return Av1BottomRightTopLeftConstants.HasTopRight(partition, blockSize, thisBlockIndex); - } - } - /// /// Prepares normative reference-edge samples and runs the selected intra predictor. /// diff --git a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1IntraSuperblockEncoderTests.cs b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1IntraSuperblockEncoderTests.cs index 9e426f5a52..98d39d06e5 100644 --- a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1IntraSuperblockEncoderTests.cs +++ b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1IntraSuperblockEncoderTests.cs @@ -613,6 +613,12 @@ public class Av1IntraSuperblockEncoderTests [InlineData((int)Av1PredictionMode.Paeth)] [InlineData((int)Av1PredictionMode.SmoothVertical)] [InlineData((int)Av1PredictionMode.SmoothHorizontal)] + [InlineData((int)Av1PredictionMode.Directional135Degrees)] + [InlineData((int)Av1PredictionMode.Directional203Degrees)] + [InlineData((int)Av1PredictionMode.Directional157Degrees)] + [InlineData((int)Av1PredictionMode.Directional67Degrees)] + [InlineData((int)Av1PredictionMode.Directional113Degrees)] + [InlineData((int)Av1PredictionMode.Directional45Degrees)] public void ProductionTileSelectsModeFromCurrentReconstruction(int expectedModeValue) { const int Width = 16; @@ -621,10 +627,42 @@ public class Av1IntraSuperblockEncoderTests const byte LeftReference = 208; const int QIndex = 1; Av1PredictionMode expectedMode = (Av1PredictionMode)expectedModeValue; + bool isDiagonal = expectedMode is >= Av1PredictionMode.Directional45Degrees and <= Av1PredictionMode.Directional67Degrees; int cornerReference = expectedMode == Av1PredictionMode.Horizontal ? LeftReference : expectedMode == Av1PredictionMode.Vertical ? TopReference : 128; + Span directionalTarget = stackalloc byte[64]; + if (isDiagonal) + { + Span aboveStorage = stackalloc byte[17]; + Span above = aboveStorage[1..]; + Span leftStorage = stackalloc byte[17]; + Span left = leftStorage[1..]; + aboveStorage[0] = 128; + leftStorage[0] = 128; + for (int i = 0; i < 8; i++) + { + above[i] = (byte)(32 + (i * 24)); + left[i] = (byte)(224 - (i * 24)); + } + + above[8..].Fill(above[7]); + left[8..].Fill(left[7]); + + // Directional arithmetic has separate byte-exact reference coverage. This fixture uses its scalar + // path only to isolate production mode traversal, reference gathering, and rate-distortion selection. + Av1DirectionalIntraPredictor.PredictScalar( + directionalTarget, + 8, + Av1TransformSize.Size8x8, + above, + left, + false, + false, + expectedMode.ToAngle()); + } + ObuColorConfig colorConfig = new() { IsMonochrome = true, @@ -667,11 +705,19 @@ public class Av1IntraSuperblockEncoderTests { value = x < 8 ? cornerReference - : expectedMode == Av1PredictionMode.Paeth ? 40 + (columnIndex * 20) : TopReference; + : isDiagonal + ? 32 + (columnIndex * 24) + : expectedMode == Av1PredictionMode.Paeth ? 40 + (columnIndex * 20) : TopReference; } else if (x < 8) { - value = expectedMode == Av1PredictionMode.Paeth ? 200 - (rowIndex * 20) : LeftReference; + value = isDiagonal + ? 224 - (rowIndex * 24) + : expectedMode == Av1PredictionMode.Paeth ? 200 - (rowIndex * 20) : LeftReference; + } + else if (isDiagonal) + { + value = directionalTarget[(rowIndex * 8) + columnIndex]; } else if (expectedMode == Av1PredictionMode.Paeth) { @@ -740,6 +786,131 @@ public class Av1IntraSuperblockEncoderTests Assert.NotEqual(0, tileWriter.GetTileData(0).Length); } + [Fact] + public void ProductionDirectionalModesConsumeAvailableExtendedEdges() + { + const int Width = 72; + const int Height = 16; + const int QIndex = 1; + ObuColorConfig colorConfig = new() + { + IsMonochrome = true, + SubSamplingX = true, + SubSamplingY = true, + BitDepth = Av1BitDepth.EightBit + }; + + using Av1EncoderFrameBuffer source = new( + Configuration.Default, + Width, + Height, + 8, + Av1ColorFormat.Yuv400, + 0, + 0); + + using Av1EncoderFrameBuffer reconstruction = new( + Configuration.Default, + Width, + Height, + 8, + Av1ColorFormat.Yuv400, + 0, + 0); + + Buffer2DRegion sourcePlane = source.Frame.CodedView.GetPlane(Av1Plane.Y); + FillPlane(sourcePlane, (byte)128); + + Span aboveStorage = stackalloc byte[17]; + Span above = aboveStorage[1..]; + Span leftStorage = stackalloc byte[17]; + Span left = leftStorage[1..]; + aboveStorage[0] = 128; + leftStorage[0] = 128; + for (int i = 0; i < 16; i++) + { + above[i] = (byte)(32 + (i * 12)); + left[i] = (byte)(224 - (i * 12)); + } + + Span topRightTarget = stackalloc byte[64]; + Span bottomLeftTarget = stackalloc byte[64]; + Span predictionScratch = stackalloc byte[64]; + Av1DirectionalIntraPredictor.Predict( + topRightTarget, + 8, + Av1TransformSize.Size8x8, + above, + left, + false, + false, + 45, + predictionScratch); + + Av1DirectionalIntraPredictor.Predict( + bottomLeftTarget, + 8, + Av1TransformSize.Size8x8, + above, + left, + false, + false, + 203, + predictionScratch); + + // The lower-left target consumes top-right samples from the already reconstructed row above. + // The upper-right superblock target consumes bottom-left samples from the completed superblock to its left. + for (int y = 0; y < Height; y++) + { + Span row = sourcePlane.DangerousGetRowSpan(y); + if (y < 8) + { + above.CopyTo(row[..16]); + bottomLeftTarget.Slice(y * 8, 8).CopyTo(row.Slice(64, 8)); + } + else + { + topRightTarget.Slice((y - 8) * 8, 8).CopyTo(row[..8]); + } + + row.Slice(56, 8).Fill(left[y]); + } + + ClearPlane(reconstruction.Luma); + using Av1EncoderModeInfoBuffer modeInfo = new(Configuration.Default, Width, Height, disallow4x4AllFrames: true); + Av1PictureControlSet pictureTemplate = CreatePicture(modeInfo, colorConfig, use128x128Superblock: false, QIndex); + using Av1EncoderPictureBuffer picture = new( + Configuration.Default, + pictureTemplate.Sequence.SequenceHeader, + pictureTemplate.Parent.FrameHeader, + Width, + Height); + + using Av1EncoderCoefficientBuffer coefficients = new( + Configuration.Default, + pictureTemplate.Sequence.SequenceHeader, + Width, + Height); + + using Av1EncoderSuperblockWorkspace superblockWorkspace = new(Configuration.Default); + using Av1EncoderBlockWorkspace blockWorkspace = new(Configuration.Default); + using Av1IntraTileWriter tileWriter = new( + Configuration.Default, + source.Frame, + reconstruction.Frame, + picture.Picture, + coefficients, + superblockWorkspace, + blockWorkspace, + initialSize: 2048); + + ref Av1MacroBlockModeInfo topRightBlock = ref picture.Picture.GetMacroBlockModeInfo(new Point(0, 2)); + ref Av1MacroBlockModeInfo bottomLeftBlock = ref picture.Picture.GetMacroBlockModeInfo(new Point(16, 0)); + Assert.Equal(Av1PredictionMode.Directional45Degrees, topRightBlock.Block.Mode); + Assert.Equal(Av1PredictionMode.Directional203Degrees, bottomLeftBlock.Block.Mode); + Assert.NotEqual(0, tileWriter.GetTileData(0).Length); + } + [Fact] public void TileWriterMapsClippedRasterTraversalToEverySuperblockCoefficientSegment() {