Browse Source

Own aligned AV1 mode information

pull/2633/head
James Jackson-South 1 month ago
parent
commit
a5c095fbe7
  1. 2
      HEIF_IMPLEMENTATION_PLAN.md
  2. 92
      src/ImageSharp/Formats/Heif/Av1/Tiling/Av1EncoderModeInfoBuffer.cs
  3. 131
      tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderModeInfoBufferTests.cs

2
HEIF_IMPLEMENTATION_PLAN.md

@ -829,7 +829,7 @@ Encoder verification contract:
- [~] Current-libaom `av1_quantize_fp_no_qmatrix` arithmetic is implemented as a closed generic forward-quantizer family with Vector512, Vector256, Vector128, and scalar paths, raster-order output, coded 64-point coefficient limits, and scan-order EOB selection. Transform search, coefficient optimization, and lossless behavior remain. - [~] Current-libaom `av1_quantize_fp_no_qmatrix` arithmetic is implemented as a closed generic forward-quantizer family with Vector512, Vector256, Vector128, and scalar paths, raster-order output, coded 64-point coefficient limits, and scan-order EOB selection. Transform search, coefficient optimization, and lossless behavior remain.
- [ ] Implement real rate-distortion selection and make quality and effort change work, size, and output quality. - [ ] Implement real rate-distortion selection and make quality and effort change work, size, and output quality.
- [~] The tile writer now publishes one packed coefficient context per covered 4x4 edge unit and derives luma/chroma skip plus DC-sign contexts from the complete transform edges using current-libaom units. Complete tile traversal, initialized picture state, and verified CDF update behavior remain. - [~] The tile writer now publishes one packed coefficient context per covered 4x4 edge unit and derives luma/chroma skip plus DC-sign contexts from the complete transform edges using current-libaom units. Complete tile traversal, initialized picture state, and verified CDF update behavior remain.
- [~] Encoder mode information now uses a frame-owned integer alias grid over a packed 8-byte value allocation, matching current libaom's `mi_grid_base` and `mi_alloc` relationship without a managed object or reference per 4x4 entry. Every coded 4x4 cell covered by square, rectangular, or clipped edge blocks maps to its owning allocation entry before context-dependent symbols are written. At 4K, mode values occupy about 4.0 MiB and the alias grid about 2.0 MiB. Packed syntax, relative neighbor lookup, full block mapping, writer traversal, entropy, and OBU coverage pass 1,947 of 1,947 direct net11 VSTest cases in Release; complete mode decision still remains. - [~] Encoder mode information now uses a frame-owned integer alias grid over a packed 8-byte value allocation, matching current libaom's `mi_grid_base` and `mi_alloc` relationship without a managed object or reference per 4x4 entry. The visible dimensions are aligned to eight luma samples, the grid stride and allocated row count are aligned to 32 mode-information units, and optional 8x8 allocation granularity reduces the value store in both dimensions exactly as current libaom does. One clean ImageSharp byte owner contains both independently typed regions, reducing libaom's two allocation lifetimes to one without a copy. At 4K, the 4x4 layout occupies about 6.0 MiB in total; the 8x8 layout occupies about 3.0 MiB. Exact geometry, clean allocation, typed lengths, aligned mapping, untouched row padding, and exactly-once return pass 4 of 4 direct net11 VSTest cases in Release. Every coded 4x4 cell covered by square, rectangular, or clipped edge blocks maps to its owning allocation entry before context-dependent symbols are written. Packed syntax, relative neighbor lookup, full block mapping, writer traversal, entropy, and OBU coverage pass 1,947 of 1,947 direct net11 VSTest cases in Release; complete mode decision still remains.
- [~] The final-block decision workspace uses one reusable 10.3 KiB ImageSharp allocator owner. It contains 1,024 explicitly packed 10-byte final-block entries and the 341 preorder partition bytes required by a complete 128x128-through-8x8 quadtree, replacing separate managed arrays. Construction and the explicit per-superblock reset initialize every syntax field, including the nonzero sentinel that disables filter-intra prediction; pooled palette, quantizer, prediction, and partition bytes cannot leak into the next decision pass. Exact allocation, size, initialization, reset, return, repeated-run, writer, entropy, and OBU coverage pass 1,957 of 1,957 direct net11 VSTest cases in Release; complete mode decision still remains. - [~] The final-block decision workspace uses one reusable 10.3 KiB ImageSharp allocator owner. It contains 1,024 explicitly packed 10-byte final-block entries and the 341 preorder partition bytes required by a complete 128x128-through-8x8 quadtree, replacing separate managed arrays. Construction and the explicit per-superblock reset initialize every syntax field, including the nonzero sentinel that disables filter-intra prediction; pooled palette, quantizer, prediction, and partition bytes cannot leak into the next decision pass. Exact allocation, size, initialization, reset, return, repeated-run, writer, entropy, and OBU coverage pass 1,957 of 1,957 direct net11 VSTest cases in Release; complete mode decision still remains.
- [~] Finalized transform coefficients and packed EOB/type state now use raster-ordered, per-superblock plane segments matching current libaom's coefficient-pool geometry. One ImageSharp allocator owner replaces libaom's separate coefficient, EOB, and entropy-context allocations while preserving the full 1024 luma and 256-per-chroma 4x4 state capacity of a 128x128 4:2:0 superblock. The transform-block boundary can populate the owner's quantized coefficient and state slices directly; mode-decision traversal still needs to select and invoke it. - [~] Finalized transform coefficients and packed EOB/type state now use raster-ordered, per-superblock plane segments matching current libaom's coefficient-pool geometry. One ImageSharp allocator owner replaces libaom's separate coefficient, EOB, and entropy-context allocations while preserving the full 1024 luma and 256-per-chroma 4x4 state capacity of a 128x128 4:2:0 superblock. The transform-block boundary can populate the owner's quantized coefficient and state slices directly; mode-decision traversal still needs to select and invoke it.
- [~] Tile partition writing now follows current libaom's recursive `write_modes_sb` preorder traversal and `update_ext_partition_context` edge updates directly. Bottom-edge blocks use the horizontal-alike partition CDF and right-edge blocks use the vertical-alike CDF; byte-exact regressions cover both paths after the previous calls were found reversed. Lossless chroma-from-luma availability now uses the subsampled plane block size shared with the decoder instead of the lossy 32x32 limit, preserving the correct UV-mode alphabet for each segment. The obsolete SVT-derived global geometry catalog and its unimplemented lookup are removed; transform geometry is derived in libaom's bounded 64x64 residual order, fixed intra transform-size symbols use the reference depth and neighbor contexts, and each derived transform size is persisted to the frame-owned mode information before the entropy snapshot and coefficient traversal consume it. Frame-edge and segmentation syntax use mode-information units, and 128x128 CDEF units use libaom's 0-to-3 indexing and first-block strength ownership. The focused transform-state regression passes 3 of 3 direct net11 VSTest cases in Release. Writer, entropy, and OBU coverage passes 1,957 of 1,957 direct net11 VSTest cases in Release, with 20 of 20 focused encoder and decoder chroma-from-luma cases. Partition and mode analysis still need to populate these retained decisions; variable inter-transform syntax remains part of later inter-frame support. - [~] Tile partition writing now follows current libaom's recursive `write_modes_sb` preorder traversal and `update_ext_partition_context` edge updates directly. Bottom-edge blocks use the horizontal-alike partition CDF and right-edge blocks use the vertical-alike CDF; byte-exact regressions cover both paths after the previous calls were found reversed. Lossless chroma-from-luma availability now uses the subsampled plane block size shared with the decoder instead of the lossy 32x32 limit, preserving the correct UV-mode alphabet for each segment. The obsolete SVT-derived global geometry catalog and its unimplemented lookup are removed; transform geometry is derived in libaom's bounded 64x64 residual order, fixed intra transform-size symbols use the reference depth and neighbor contexts, and each derived transform size is persisted to the frame-owned mode information before the entropy snapshot and coefficient traversal consume it. Frame-edge and segmentation syntax use mode-information units, and 128x128 CDEF units use libaom's 0-to-3 indexing and first-block strength ownership. The focused transform-state regression passes 3 of 3 direct net11 VSTest cases in Release. Writer, entropy, and OBU coverage passes 1,957 of 1,957 direct net11 VSTest cases in Release, with 20 of 20 focused encoder and decoder chroma-from-luma cases. Partition and mode analysis still need to populate these retained decisions; variable inter-transform syntax remains part of later inter-frame support.

92
src/ImageSharp/Formats/Heif/Av1/Tiling/Av1EncoderModeInfoBuffer.cs

@ -0,0 +1,92 @@
// Copyright (c) Six Labors.
// Licensed under the Six Labors Split License.
using System.Buffers;
using System.Runtime.CompilerServices;
using SixLabors.ImageSharp.Memory;
namespace SixLabors.ImageSharp.Formats.Heif.Av1.Tiling;
/// <summary>
/// Owns the allocation-index grid and mode-information values for one encoded AV1 frame.
/// </summary>
internal sealed class Av1EncoderModeInfoBuffer : IDisposable
{
private const int CodedDimensionAlignmentLog2 = 3;
private const int ModeInfoAlignmentLog2 = Av1Constants.MaxSuperBlockSizeLog2 - Av1Constants.ModeInfoSizeLog2;
private IMemoryOwner<byte>? owner;
private readonly ByteMemoryManager<int> grid;
private readonly ByteMemoryManager<Av1MacroBlockModeInfo> allocation;
/// <summary>
/// Initializes a new instance of the <see cref="Av1EncoderModeInfoBuffer"/> class.
/// </summary>
/// <param name="configuration">The configuration providing the memory allocator.</param>
/// <param name="frameWidth">The visible frame width in luma samples.</param>
/// <param name="frameHeight">The visible frame height in luma samples.</param>
/// <param name="disallow4x4AllFrames">Whether each allocated mode-information value represents an 8x8 region.</param>
public Av1EncoderModeInfoBuffer(
Configuration configuration,
int frameWidth,
int frameHeight,
bool disallow4x4AllFrames)
{
this.ModeInfoColumnCount = Av1Math.AlignPowerOf2(frameWidth, CodedDimensionAlignmentLog2) >> Av1Constants.ModeInfoSizeLog2;
this.ModeInfoRowCount = Av1Math.AlignPowerOf2(frameHeight, CodedDimensionAlignmentLog2) >> Av1Constants.ModeInfoSizeLog2;
this.ModeInfoStride = Av1Math.AlignPowerOf2(this.ModeInfoColumnCount, ModeInfoAlignmentLog2);
int alignedModeInfoRowCount = Av1Math.AlignPowerOf2(this.ModeInfoRowCount, ModeInfoAlignmentLog2);
int allocationShift = disallow4x4AllFrames ? 1 : 0;
int gridLength = checked(this.ModeInfoStride * alignedModeInfoRowCount);
int allocationLength = checked((this.ModeInfoStride >> allocationShift) * (alignedModeInfoRowCount >> allocationShift));
int gridByteLength = checked(gridLength * sizeof(int));
int allocationByteLength = checked(allocationLength * Unsafe.SizeOf<Av1MacroBlockModeInfo>());
int storageLength = checked(gridByteLength + allocationByteLength);
// The pointer grid and value allocation share one frame lifetime. Packing both regions into one clean
// owner retains libaom's independent typed layouts without its separate allocation and cleanup paths.
this.owner = configuration.MemoryAllocator.Allocate<byte>(storageLength, AllocationOptions.Clean);
Memory<byte> storage = this.owner.Memory[..storageLength];
this.grid = new ByteMemoryManager<int>(storage[..gridByteLength]);
this.allocation = new ByteMemoryManager<Av1MacroBlockModeInfo>(storage.Slice(gridByteLength, allocationByteLength));
this.Disallow4x4AllFrames = disallow4x4AllFrames;
}
/// <summary>
/// Gets the visible frame height in 4x4 mode-information units.
/// </summary>
public int ModeInfoRowCount { get; }
/// <summary>
/// Gets the visible frame width in 4x4 mode-information units.
/// </summary>
public int ModeInfoColumnCount { get; }
/// <summary>
/// Gets the aligned row stride of the allocation-index grid in 4x4 mode-information units.
/// </summary>
public int ModeInfoStride { get; }
/// <summary>
/// Gets a value indicating whether each allocated mode-information value represents an 8x8 region.
/// </summary>
public bool Disallow4x4AllFrames { get; }
/// <summary>
/// Gets the frame grid that maps each 4x4 position to its mode-information allocation index.
/// </summary>
public Memory<int> Grid => this.grid.Memory;
/// <summary>
/// Gets the contiguous mode-information values addressed by <see cref="Grid"/>.
/// </summary>
public Memory<Av1MacroBlockModeInfo> Allocation => this.allocation.Memory;
/// <summary>
/// Returns the packed frame storage to the configured memory allocator.
/// </summary>
public void Dispose()
{
this.owner?.Dispose();
this.owner = null;
}
}

131
tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderModeInfoBufferTests.cs

@ -0,0 +1,131 @@
// Copyright (c) Six Labors.
// Licensed under the Six Labors Split License.
using SixLabors.ImageSharp.Formats.Heif.Av1;
using SixLabors.ImageSharp.Formats.Heif.Av1.OpenBitstreamUnit;
using SixLabors.ImageSharp.Formats.Heif.Av1.Prediction;
using SixLabors.ImageSharp.Formats.Heif.Av1.Tiling;
using SixLabors.ImageSharp.Memory;
using SixLabors.ImageSharp.Tests.Memory;
namespace SixLabors.ImageSharp.Tests.Formats.Heif.Av1;
public class Av1EncoderModeInfoBufferTests
{
[Theory]
[InlineData(false, 1024, 1024, 12_288)]
[InlineData(true, 1024, 256, 6_144)]
public void ConstructorMatchesLibaomAlignedModeInfoGeometry(
bool disallow4x4,
int expectedGridLength,
int expectedAllocationLength,
int expectedStorageLength)
{
TestMemoryAllocator allocator = new();
allocator.EnableNonThreadSafeLogging();
Configuration configuration = Configuration.Default.Clone();
configuration.MemoryAllocator = allocator;
TestMemoryAllocator.AllocationRequest allocation;
using (Av1EncoderModeInfoBuffer buffer = new(configuration, 65, 33, disallow4x4))
{
allocation = Assert.Single(allocator.AllocationLog);
Assert.Empty(allocator.ReturnLog);
Assert.Equal(typeof(byte), allocation.ElementType);
Assert.Equal(AllocationOptions.Clean, allocation.AllocationOptions);
Assert.Equal(expectedStorageLength, allocation.Length);
Assert.Equal(18, buffer.ModeInfoColumnCount);
Assert.Equal(10, buffer.ModeInfoRowCount);
Assert.Equal(32, buffer.ModeInfoStride);
Assert.Equal(disallow4x4, buffer.Disallow4x4AllFrames);
Assert.Equal(expectedGridLength, buffer.Grid.Length);
Assert.Equal(expectedAllocationLength, buffer.Allocation.Length);
}
TestMemoryAllocator.ReturnRequest returned = Assert.Single(allocator.ReturnLog);
Assert.Equal(allocation.AllocationId, returned.AllocationId);
}
[Theory]
[InlineData(false, 2, 3, 98)]
[InlineData(true, 2, 2, 17)]
public void PictureMappingUsesPackedAlignedStorage(
bool disallow4x4,
int column,
int row,
int expectedAllocationOffset)
{
using Av1EncoderModeInfoBuffer buffer = new(Configuration.Default, 65, 33, disallow4x4);
Av1PictureControlSet picture = CreatePicture(buffer);
Point position = new(column, row);
ref Av1MacroBlockModeInfo modeInfo = ref picture.GetMacroBlockModeInfo(position);
modeInfo.Block.Mode = Av1PredictionMode.Paeth;
picture.MapModeInfoBlock(position, Av1BlockSize.Block8x8);
Assert.Equal(Av1PredictionMode.Paeth, buffer.Allocation.Span[expectedAllocationOffset].Block.Mode);
int alignedRowCount = buffer.Grid.Length / buffer.ModeInfoStride;
for (int y = 0; y < alignedRowCount; y++)
{
for (int x = 0; x < buffer.ModeInfoStride; x++)
{
int gridOffset = (y * buffer.ModeInfoStride) + x;
bool isMapped = y >= row && y < row + 2 && x >= column && x < column + 2;
Assert.Equal(isMapped ? expectedAllocationOffset : 0, buffer.Grid.Span[gridOffset]);
if (isMapped)
{
Assert.Equal(Av1PredictionMode.Paeth, picture.GetFromModeInfoGrid(new Point(x, y)).Block.Mode);
}
}
}
}
private static Av1PictureControlSet CreatePicture(Av1EncoderModeInfoBuffer buffer)
{
ObuTileGroupHeader tiles = new()
{
TileColumnCount = 1,
TileRowCount = 1
};
tiles.TileColumnStartModeInfo[1] = buffer.ModeInfoColumnCount;
tiles.TileRowStartModeInfo[1] = buffer.ModeInfoRowCount;
ObuSequenceHeader sequenceHeader = new();
ObuFrameHeader frameHeader = new()
{
ModeInfoColumnCount = buffer.ModeInfoColumnCount,
ModeInfoRowCount = buffer.ModeInfoRowCount,
TilesInfo = tiles
};
return new Av1PictureControlSet
{
PartitionContexts = [],
LuminanceDcSignLevelCoefficientNeighbors = [],
CrDcSignLevelCoefficientNeighbors = [],
CbDcSignLevelCoefficientNeighbors = [],
TransformFunctionContexts = [],
Sequence = new Av1SequenceControlSet { SequenceHeader = sequenceHeader },
Parent = new Av1PictureParentControlSet
{
Common = new Av1EncoderCommon
{
ModeInfoColumnCount = buffer.ModeInfoColumnCount,
ModeInfoRowCount = buffer.ModeInfoRowCount,
ModeInfoStride = buffer.ModeInfoStride,
FrameSize = new ObuFrameSize(),
TilesInfo = tiles
},
FrameHeader = frameHeader,
PreviousQIndex = []
},
SegmentationNeighborMap = [],
ModeInfoGrid = buffer.Grid,
ModeInfoAllocation = buffer.Allocation,
ModeInfoStride = buffer.ModeInfoStride,
Disallow4x4AllFrames = buffer.Disallow4x4AllFrames,
CdefPreset = []
};
}
}
Loading…
Cancel
Save