Browse Source

Compose AV1 intra block encoding

pull/2633/head
James Jackson-South 1 month ago
parent
commit
7cb6292040
  1. 6
      HEIF_IMPLEMENTATION_PLAN.md
  2. 176
      src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1TransformBlockEncoder.cs
  3. 359
      tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs

6
HEIF_IMPLEMENTATION_PLAN.md

@ -819,7 +819,7 @@ Encoder verification contract:
### 6. Build the complete AV1 frame encoder
- [~] SIMD-first RGB-to-native-plane conversion now feeds eight-bit and high-bit-depth bordered AV1 source frames directly, preserving ImageSharp's arbitrary packed-pixel input contract without an intermediate full-frame native-plane copy.
- [~] Forward transform families, transform workspace, and a zero-allocation transform-then-quantize block boundary exist locally. The block boundary matches current libaom's selected full-transform, quantization, qcoeff, dqcoeff, EOB, and transform-type data flow. One reusable 61 KiB allocator owner supplies tightly packed residual, aligned transform-coefficient, dequantized-coefficient, and transform scratch spans across transform blocks; quantized coefficients write directly to the retained frame coefficient owner instead of being duplicated. The exact owner type, length, span partitioning, and exactly-once return are verified, and the adjacent residual, transform, quantizer, and block tests pass 12 of 12 through direct net11 VSTest in Release. Frame traversal still needs to select blocks and supply coefficient-owner slices.
- [~] Forward transform families, transform workspace, and an allocation-free DC intra block boundary exist locally. For eight-bit and high-bit-depth samples, the composed boundary now follows current libaom's encoder order: predict into the reconstruction plane, subtract prediction from source, transform, quantize into separate qcoeff and dqcoeff storage, retain EOB and transform type, and inverse-transform only when EOB is nonzero so later blocks consume decoder-identical references. Prediction and subtraction retain their SIMD-first operators, independent source and reconstruction strides are preserved, and no frame-sized or per-block buffer is introduced. One reusable 61 KiB allocator owner supplies tightly packed residual, aligned transform-coefficient, dequantized-coefficient, and transform scratch spans across transform blocks; quantized coefficients write directly to the retained frame coefficient owner instead of being duplicated. Stage-by-stage scalar-oracle, padding, retained-syntax, and zero-allocation coverage passes 6 of 6 through direct net11 VSTest in Release. Frame traversal still needs to select blocks, gather contiguous left references, and supply coefficient-owner slices.
- [~] Symbol writer, coefficient writer, and tile writer fragments exist locally.
- [~] A non-owning encoder-frame view now separates visible conversion regions from coded regions and performs complete left, top, right, bottom, and corner extension across each bordered plane. Current libaom uses 8-sample-aligned coded dimensions, a 32-sample-aligned luma stride with chroma stride derived from it, and a 64-pixel luma border for non-resized all-intra encoding. One operation-ready frame owner now rents the aligned Y, U, and V storage contiguously, exposes non-owning `Buffer2D` plane views, and returns the rent exactly once. A 4K 4:2:0 frame occupies about 13.0 MiB at 8-bit or 26.0 MiB at 10/12-bit; source and reconstruction therefore remain distinct frame owners rather than adding a full-frame copy. The corrected tests use this real ownership path and verify the exact 54 KiB 64x64 4:2:0 rent. The frame-encoder boundary converts packed pixels directly into the source owner before extension; the containing encode operation still needs to instantiate matching source and reconstruction owners with ordinary `using` lifetimes.
- [~] Temporal delimiter, sequence header, frame header, and combined-frame tile-group writing exist locally. The remaining required metadata, padding, and encoder-wide syntax paths are not complete.
@ -831,11 +831,11 @@ Encoder verification contract:
- [~] The tile writer now publishes one packed coefficient context per covered 4x4 edge unit and derives luma/chroma skip plus DC-sign contexts from the complete transform edges using current-libaom units. Partition, transform, and coefficient neighbor state now retains only the above and left context regions used by current libaom; the unused third top-left region, its granularity state, and its unused sentinel are removed. One clean allocation contains the two active edges, and the exact requested length plus exactly-once return pass with the complete 77-case coefficient and entropy class in direct net11 VSTest Release. Complete tile traversal, initialized picture state, and verified CDF update behavior remain.
- [~] Encoder mode information now uses a frame-owned integer alias grid over a packed 8-byte value allocation, matching current libaom's `mi_grid_base` and `mi_alloc` relationship without a managed object or reference per 4x4 entry. The visible dimensions are aligned to eight luma samples, the grid stride and allocated row count are aligned to 32 mode-information units, and optional 8x8 allocation granularity reduces the value store in both dimensions exactly as current libaom does. One clean ImageSharp byte owner contains both independently typed regions, reducing libaom's two allocation lifetimes to one without a copy. At 4K, the 4x4 layout occupies about 6.0 MiB in total; the 8x8 layout occupies about 3.0 MiB. Exact geometry, clean allocation, typed lengths, aligned mapping, untouched row padding, and exactly-once return pass 4 of 4 direct net11 VSTest cases in Release. Every coded 4x4 cell covered by square, rectangular, or clipped edge blocks maps to its owning allocation entry before context-dependent symbols are written. Packed syntax, relative neighbor lookup, full block mapping, writer traversal, entropy, and OBU coverage pass 1,947 of 1,947 direct net11 VSTest cases in Release; complete mode decision still remains.
- [~] The final-block decision workspace uses one reusable 10.3 KiB ImageSharp allocator owner. It contains 1,024 explicitly packed 10-byte final-block entries and the 341 preorder partition bytes required by a complete 128x128-through-8x8 quadtree, replacing separate managed arrays. Construction and the explicit per-superblock reset initialize every syntax field, including the nonzero sentinel that disables filter-intra prediction; pooled palette, quantizer, prediction, and partition bytes cannot leak into the next decision pass. Exact allocation, size, initialization, reset, return, repeated-run, writer, entropy, and OBU coverage pass 1,957 of 1,957 direct net11 VSTest cases in Release; complete mode decision still remains.
- [~] Finalized transform coefficients and packed EOB/type state now use raster-ordered, per-superblock plane segments matching current libaom's coefficient-pool geometry. One ImageSharp allocator owner replaces libaom's separate coefficient, EOB, and entropy-context allocations while preserving the full 1024 luma and 256-per-chroma 4x4 state capacity of a 128x128 4:2:0 superblock. The transform-block boundary can populate the owner's quantized coefficient and state slices directly; mode-decision traversal still needs to select and invoke it.
- [~] Finalized transform coefficients and packed EOB/type state now use raster-ordered, per-superblock plane segments matching current libaom's coefficient-pool geometry. One ImageSharp allocator owner replaces libaom's separate coefficient, EOB, and entropy-context allocations while preserving the full 1024 luma and 256-per-chroma 4x4 state capacity of a 128x128 4:2:0 superblock. The composed intra block boundary can populate the owner's quantized coefficient and state slices while updating the caller-owned reconstruction plane directly; mode-decision traversal still needs to select and invoke it.
- [~] Tile partition writing now follows current libaom's recursive `write_modes_sb` preorder traversal and `update_ext_partition_context` edge updates directly. Bottom-edge blocks use the horizontal-alike partition CDF and right-edge blocks use the vertical-alike CDF; byte-exact regressions cover both paths after the previous calls were found reversed. Lossless chroma-from-luma availability now uses the subsampled plane block size shared with the decoder instead of the lossy 32x32 limit, preserving the correct UV-mode alphabet for each segment. The obsolete SVT-derived global geometry catalog and its unimplemented lookup are removed; transform geometry is derived in libaom's bounded 64x64 residual order, fixed intra transform-size symbols use the reference depth and neighbor contexts, and each derived transform size is persisted to the frame-owned mode information before the entropy snapshot and coefficient traversal consume it. Frame-edge and segmentation syntax use mode-information units, and 128x128 CDEF units use libaom's 0-to-3 indexing and first-block strength ownership. The focused transform-state regression passes 3 of 3 direct net11 VSTest cases in Release. Writer, entropy, and OBU coverage passes 1,957 of 1,957 direct net11 VSTest cases in Release, with 20 of 20 focused encoder and decoder chroma-from-luma cases. Partition and mode analysis still need to populate these retained decisions; variable inter-transform syntax remains part of later inter-frame support.
- [ ] Implement legal deblocking, CDEF, restoration, super-resolution, and film-grain signaling decisions.
- [~] The coefficient symbol encoder now reuses tile-lifetime level and context workspaces instead of allocating per transform, defers both coefficient rents until the first nonzero transform block, and disposes all tile scratch independently from the detached encoded bytes. Its range coder matches current libaom's 64-bit coding window, bulk big-endian byte flush, and backward carry propagation while using one byte of allocator scratch per estimated output byte instead of the former 16-bit pre-carry storage. The reference-type symbol encoder is passed normally through tile traversal, and the operation boundary owns the allocator-backed item payload stream for exactly one synchronous encode. Every remaining encoder fragment must be audited before it becomes active.
- [~] The planar conversion, residual construction, forward transform, and forward quantizer use descending SIMD dispatch: Vector512, Vector256, Vector128, then scalar. Residual construction matches current libaom's exact source-minus-prediction arithmetic for 8-bit and high-bit-depth planes, preserves independent row strides and unaligned starts, and writes directly into caller-owned signed-short storage without allocation. Apply the same rule to every later hot-path family.
- [~] The planar conversion, DC intra prediction, residual construction, forward transform, and forward quantizer use descending SIMD dispatch: Vector512, Vector256, Vector128, then scalar. Residual construction matches current libaom's exact source-minus-prediction arithmetic for 8-bit and high-bit-depth planes, preserves independent row strides and unaligned starts, and writes directly into caller-owned signed-short storage without allocation. The composed block path delegates arithmetic to those closed operators and adds no allocation. Apply the same rule to every later hot-path family.
- [~] Residual tests verify misaligned planes, independent source, prediction, and destination strides, SIMD remainders, untouched padding, 8-bit, 10-bit, and 12-bit precision, every operator width independently of host acceleration, the scalar fallback, and zero per-transform allocations.
- [~] The unused coefficient-shape transform facade and its unimplemented N2, N4, and DC-only branches are removed. Finalized block encoding now follows the complete-transform path that current libaom uses before fast quantization; later rate-distortion search may add proven coefficient optimization without exposing inactive runtime throws.
- [~] Forward-quantizer FeatureTestRunner and zero-allocation tests compare every hardware tier with an independent scan-order scalar oracle shaped from current-main libaom. Both passed direct net11 VSTest in Release.

176
src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1TransformBlockEncoder.cs

@ -1,21 +1,193 @@
// Copyright (c) Six Labors.
// Licensed under the Six Labors Split License.
using System.Runtime.InteropServices;
using SixLabors.ImageSharp.Formats.Heif.Av1.Pipeline.Quantizers;
using SixLabors.ImageSharp.Formats.Heif.Av1.Prediction;
using SixLabors.ImageSharp.Formats.Heif.Av1.Tiling;
using SixLabors.ImageSharp.Formats.Heif.Av1.Transform;
namespace SixLabors.ImageSharp.Formats.Heif.Av1.Pipeline;
/// <summary>
/// Transforms and quantizes finalized AV1 residual blocks.
/// Predicts, transforms, quantizes, and reconstructs finalized AV1 blocks.
/// </summary>
internal static class Av1TransformBlockEncoder
{
/// <summary>
/// Encodes and reconstructs one eight-bit lossy DC intra block.
/// </summary>
/// <param name="workspace">The reusable residual, coefficient, and transform storage.</param>
/// <param name="source">The source samples.</param>
/// <param name="sourceStride">The number of source samples between rows.</param>
/// <param name="reconstruction">The reconstructed frame samples and prediction destination.</param>
/// <param name="reconstructionStride">The number of reconstruction samples between rows.</param>
/// <param name="above">The contiguous top reference samples.</param>
/// <param name="left">The contiguous left reference samples.</param>
/// <param name="hasLeft">Whether the left reference is available.</param>
/// <param name="hasAbove">Whether the top reference is available.</param>
/// <param name="quantizedCoefficients">The retained entropy-coding coefficients.</param>
/// <param name="transformSize">The selected transform dimensions.</param>
/// <param name="transformType">The selected compound transform type.</param>
/// <param name="qIndex">The segment quantizer index.</param>
/// <param name="dcDeltaQ">The plane DC quantizer adjustment.</param>
/// <param name="acDeltaQ">The plane AC quantizer adjustment.</param>
/// <param name="plane">The component plane containing the block.</param>
/// <param name="state">The retained transform type and end-of-block syntax.</param>
public static void EncodeIntraDcLossy(
Av1EncoderBlockWorkspace workspace,
ReadOnlySpan<byte> source,
int sourceStride,
Span<byte> reconstruction,
int reconstructionStride,
ReadOnlySpan<byte> above,
ReadOnlySpan<byte> left,
bool hasLeft,
bool hasAbove,
Span<int> quantizedCoefficients,
Av1TransformSize transformSize,
Av1TransformType transformType,
int qIndex,
int dcDeltaQ,
int acDeltaQ,
Av1Plane plane,
ref Av1EncoderTransformBlockState state)
{
int width = transformSize.GetWidth();
int height = transformSize.GetHeight();
// Prediction and subtraction stay in their SIMD-first operators while this method owns the required block-stage ordering.
Av1DcIntraPredictor.Predict(hasLeft, hasAbove, reconstruction, reconstructionStride, above, left, width, height);
Av1ResidualBuilder.Subtract(source, sourceStride, reconstruction, reconstructionStride, workspace.Residual, width, width, height);
EncodeLossy(
workspace,
quantizedCoefficients,
transformSize,
transformType,
qIndex,
dcDeltaQ,
acDeltaQ,
Av1BitDepth.EightBit,
ref state);
if (state.EndOfBlock > 0)
{
// Reconstructing the quantized result makes later predictions use exactly the samples a decoder will reproduce.
Av1InverseTransformer.Reconstruct8Bit(
workspace.DequantizedCoefficients,
reconstruction,
reconstructionStride,
transformSize,
transformType,
(int)plane,
state.EndOfBlock,
false,
workspace.TransformWorkspace);
}
}
/// <summary>
/// Encodes and reconstructs one high-bit-depth lossy DC intra block.
/// </summary>
/// <param name="workspace">The reusable residual, coefficient, and transform storage.</param>
/// <param name="source">The source samples.</param>
/// <param name="sourceStride">The number of source samples between rows.</param>
/// <param name="reconstruction">The reconstructed frame samples and prediction destination.</param>
/// <param name="reconstructionStride">The number of reconstruction samples between rows.</param>
/// <param name="above">The contiguous top reference samples.</param>
/// <param name="left">The contiguous left reference samples.</param>
/// <param name="hasLeft">Whether the left reference is available.</param>
/// <param name="hasAbove">Whether the top reference is available.</param>
/// <param name="quantizedCoefficients">The retained entropy-coding coefficients.</param>
/// <param name="transformSize">The selected transform dimensions.</param>
/// <param name="transformType">The selected compound transform type.</param>
/// <param name="qIndex">The segment quantizer index.</param>
/// <param name="dcDeltaQ">The plane DC quantizer adjustment.</param>
/// <param name="acDeltaQ">The plane AC quantizer adjustment.</param>
/// <param name="plane">The component plane containing the block.</param>
/// <param name="bitDepth">The coded sample bit depth.</param>
/// <param name="state">The retained transform type and end-of-block syntax.</param>
public static void EncodeIntraDcLossy(
Av1EncoderBlockWorkspace workspace,
ReadOnlySpan<ushort> source,
int sourceStride,
Span<ushort> reconstruction,
int reconstructionStride,
ReadOnlySpan<ushort> above,
ReadOnlySpan<ushort> left,
bool hasLeft,
bool hasAbove,
Span<int> quantizedCoefficients,
Av1TransformSize transformSize,
Av1TransformType transformType,
int qIndex,
int dcDeltaQ,
int acDeltaQ,
Av1Plane plane,
Av1BitDepth bitDepth,
ref Av1EncoderTransformBlockState state)
{
int width = transformSize.GetWidth();
int height = transformSize.GetHeight();
// Valid high-bit-depth samples remain below the sign bit, so signed transform lanes can share the unsigned frame storage.
Span<short> signedReconstruction = MemoryMarshal.Cast<ushort, short>(reconstruction);
ReadOnlySpan<short> signedAbove = MemoryMarshal.Cast<ushort, short>(above);
ReadOnlySpan<short> signedLeft = MemoryMarshal.Cast<ushort, short>(left);
Av1DcIntraPredictor.Predict(
hasLeft,
hasAbove,
signedReconstruction,
reconstructionStride,
signedAbove,
signedLeft,
width,
height,
bitDepth.GetBitCount());
Av1ResidualBuilder.Subtract(
source,
sourceStride,
reconstruction,
reconstructionStride,
workspace.Residual,
width,
width,
height);
EncodeLossy(
workspace,
quantizedCoefficients,
transformSize,
transformType,
qIndex,
dcDeltaQ,
acDeltaQ,
bitDepth,
ref state);
if (state.EndOfBlock > 0)
{
// Reconstructing the quantized result makes later predictions use exactly the samples a decoder will reproduce.
Av1InverseTransformer.ReconstructHighBitDepth(
workspace.DequantizedCoefficients,
signedReconstruction,
reconstructionStride,
transformSize,
transformType,
(int)plane,
state.EndOfBlock,
false,
bitDepth,
workspace.TransformWorkspace);
}
}
/// <summary>
/// Applies a lossy forward transform and quantization to one residual block.
/// </summary>
/// <param name="workspace">The reusable residual, transform, and reconstruction storage.</param>
/// <param name="workspace">The reusable residual, coefficient, and transform storage.</param>
/// <param name="quantizedCoefficients">The retained entropy-coding coefficients.</param>
/// <param name="transformSize">The selected transform dimensions.</param>
/// <param name="transformType">The selected compound transform type.</param>

359
tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs

@ -1,9 +1,11 @@
// Copyright (c) Six Labors.
// Licensed under the Six Labors Split License.
using System.Runtime.InteropServices;
using SixLabors.ImageSharp.Formats.Heif.Av1;
using SixLabors.ImageSharp.Formats.Heif.Av1.Pipeline;
using SixLabors.ImageSharp.Formats.Heif.Av1.Pipeline.Quantizers;
using SixLabors.ImageSharp.Formats.Heif.Av1.Prediction;
using SixLabors.ImageSharp.Formats.Heif.Av1.Tiling;
using SixLabors.ImageSharp.Formats.Heif.Av1.Transform;
using SixLabors.ImageSharp.Tests.Memory;
@ -11,7 +13,7 @@ using SixLabors.ImageSharp.Tests.Memory;
namespace SixLabors.ImageSharp.Tests.Formats.Heif.Av1;
/// <summary>
/// Verifies the finalized forward-transform and quantization block boundary.
/// Verifies the finalized prediction, transform, quantization, and reconstruction block boundary.
/// </summary>
[Trait("Format", "Avif")]
public class Av1TransformBlockEncoderTests
@ -28,6 +30,339 @@ public class Av1TransformBlockEncoderTests
ValidateBlock(Av1TransformSize.Size64x64, Av1TransformType.DctDct, Av1BitDepth.TwelveBit, 255);
}
/// <summary>
/// Verifies that the eight-bit block boundary preserves stage ordering, strides, padding, and retained syntax.
/// </summary>
[Fact]
public void EightBitIntraDcBlockEncodingMatchesStageContracts()
{
const int SourceStride = 13;
const int ReconstructionStride = 15;
Av1TransformSize transformSize = Av1TransformSize.Size8x8;
int width = transformSize.GetWidth();
int height = transformSize.GetHeight();
int coefficientCount = transformSize.GetAdjusted().GetSize2d();
byte[] source = new byte[SourceStride * height];
byte[] expectedReconstruction = new byte[ReconstructionStride * height];
byte[] actualReconstruction = new byte[ReconstructionStride * height];
byte[] above = new byte[width];
byte[] left = new byte[height];
int[] expectedQuantized = new int[coefficientCount + 7];
int[] actualQuantized = new int[coefficientCount + 7];
using Av1EncoderBlockWorkspace expectedWorkspace = new(Configuration.Default);
using Av1EncoderBlockWorkspace actualWorkspace = new(Configuration.Default);
FillSource(source, SourceStride, width, height, byte.MaxValue);
Array.Fill(expectedReconstruction, (byte)211);
Array.Fill(actualReconstruction, (byte)211);
Array.Fill(expectedQuantized, int.MinValue);
Array.Fill(actualQuantized, int.MinValue);
for (int i = 0; i < above.Length; i++)
{
above[i] = (byte)(37 + (i * 11));
}
for (int i = 0; i < left.Length; i++)
{
left[i] = (byte)(19 + (i * 13));
}
Av1DcIntraPredictor.PredictScalar(
true,
true,
expectedReconstruction,
ReconstructionStride,
above,
left,
width,
height);
for (int y = 0; y < height; y++)
{
for (int x = 0; x < width; x++)
{
expectedWorkspace.Residual[(y * width) + x] =
(short)(source[(y * SourceStride) + x] - expectedReconstruction[(y * ReconstructionStride) + x]);
}
}
Av1EncoderTransformBlockState expectedState = default;
Av1TransformBlockEncoder.EncodeLossy(
expectedWorkspace,
expectedQuantized,
transformSize,
Av1TransformType.DctDct,
73,
-1,
3,
Av1BitDepth.EightBit,
ref expectedState);
if (expectedState.EndOfBlock > 0)
{
Av1InverseTransformer.Reconstruct8Bit(
expectedWorkspace.DequantizedCoefficients,
expectedReconstruction,
ReconstructionStride,
transformSize,
Av1TransformType.DctDct,
(int)Av1Plane.Y,
expectedState.EndOfBlock,
false,
expectedWorkspace.TransformWorkspace);
}
Av1EncoderTransformBlockState actualState = default;
Av1TransformBlockEncoder.EncodeIntraDcLossy(
actualWorkspace,
source,
SourceStride,
actualReconstruction,
ReconstructionStride,
above,
left,
true,
true,
actualQuantized,
transformSize,
Av1TransformType.DctDct,
73,
-1,
3,
Av1Plane.Y,
ref actualState);
Assert.Equal(expectedReconstruction, actualReconstruction);
Assert.Equal(expectedQuantized, actualQuantized);
Assert.Equal(expectedState.EndOfBlock, actualState.EndOfBlock);
Assert.Equal(expectedState.TransformType, actualState.TransformType);
}
/// <summary>
/// Verifies that the high-bit-depth block boundary preserves stage ordering, strides, padding, and retained syntax.
/// </summary>
[Fact]
public void HighBitDepthIntraDcBlockEncodingMatchesStageContracts()
{
const int SourceStride = 19;
const int ReconstructionStride = 23;
Av1TransformSize transformSize = Av1TransformSize.Size16x8;
Av1BitDepth bitDepth = Av1BitDepth.TenBit;
int width = transformSize.GetWidth();
int height = transformSize.GetHeight();
int coefficientCount = transformSize.GetAdjusted().GetSize2d();
ushort[] source = new ushort[SourceStride * height];
ushort[] expectedReconstruction = new ushort[ReconstructionStride * height];
ushort[] actualReconstruction = new ushort[ReconstructionStride * height];
ushort[] above = new ushort[width];
ushort[] left = new ushort[height];
int[] expectedQuantized = new int[coefficientCount + 7];
int[] actualQuantized = new int[coefficientCount + 7];
using Av1EncoderBlockWorkspace expectedWorkspace = new(Configuration.Default);
using Av1EncoderBlockWorkspace actualWorkspace = new(Configuration.Default);
FillSource(source, SourceStride, width, height, (1 << bitDepth.GetBitCount()) - 1);
Array.Fill(expectedReconstruction, (ushort)777);
Array.Fill(actualReconstruction, (ushort)777);
Array.Fill(expectedQuantized, int.MinValue);
Array.Fill(actualQuantized, int.MinValue);
for (int i = 0; i < above.Length; i++)
{
above[i] = (ushort)(173 + (i * 17));
}
for (int i = 0; i < left.Length; i++)
{
left[i] = (ushort)(91 + (i * 29));
}
Span<short> signedExpectedReconstruction = MemoryMarshal.Cast<ushort, short>(expectedReconstruction.AsSpan());
Av1DcIntraPredictor.PredictScalar(
true,
false,
signedExpectedReconstruction,
ReconstructionStride,
MemoryMarshal.Cast<ushort, short>(above),
MemoryMarshal.Cast<ushort, short>(left),
width,
height,
bitDepth.GetBitCount());
for (int y = 0; y < height; y++)
{
for (int x = 0; x < width; x++)
{
expectedWorkspace.Residual[(y * width) + x] =
(short)(source[(y * SourceStride) + x] - expectedReconstruction[(y * ReconstructionStride) + x]);
}
}
Av1EncoderTransformBlockState expectedState = default;
Av1TransformBlockEncoder.EncodeLossy(
expectedWorkspace,
expectedQuantized,
transformSize,
Av1TransformType.DctDct,
117,
-2,
4,
bitDepth,
ref expectedState);
if (expectedState.EndOfBlock > 0)
{
Av1InverseTransformer.ReconstructHighBitDepth(
expectedWorkspace.DequantizedCoefficients,
signedExpectedReconstruction,
ReconstructionStride,
transformSize,
Av1TransformType.DctDct,
(int)Av1Plane.U,
expectedState.EndOfBlock,
false,
bitDepth,
expectedWorkspace.TransformWorkspace);
}
Av1EncoderTransformBlockState actualState = default;
Av1TransformBlockEncoder.EncodeIntraDcLossy(
actualWorkspace,
source,
SourceStride,
actualReconstruction,
ReconstructionStride,
above,
left,
true,
false,
actualQuantized,
transformSize,
Av1TransformType.DctDct,
117,
-2,
4,
Av1Plane.U,
bitDepth,
ref actualState);
Assert.Equal(expectedReconstruction, actualReconstruction);
Assert.Equal(expectedQuantized, actualQuantized);
Assert.Equal(expectedState.EndOfBlock, actualState.EndOfBlock);
Assert.Equal(expectedState.TransformType, actualState.TransformType);
}
/// <summary>
/// Verifies that complete eight-bit and high-bit-depth DC block encoding uses only caller-owned storage.
/// </summary>
[Fact]
public void IntraDcBlockEncodingDoesNotAllocate()
{
const int Stride = 8;
Av1TransformSize transformSize = Av1TransformSize.Size8x8;
int coefficientCount = transformSize.GetAdjusted().GetSize2d();
byte[] source8 = new byte[Stride * Stride];
byte[] reconstruction8 = new byte[Stride * Stride];
byte[] above8 = new byte[Stride];
byte[] left8 = new byte[Stride];
ushort[] source10 = new ushort[Stride * Stride];
ushort[] reconstruction10 = new ushort[Stride * Stride];
ushort[] above10 = new ushort[Stride];
ushort[] left10 = new ushort[Stride];
int[] quantized = new int[coefficientCount];
using Av1EncoderBlockWorkspace workspace = new(Configuration.Default);
FillSource(source8, Stride, Stride, Stride, byte.MaxValue);
FillSource(source10, Stride, Stride, Stride, 1023);
Array.Fill(above8, (byte)103);
Array.Fill(left8, (byte)127);
Array.Fill(above10, (ushort)503);
Array.Fill(left10, (ushort)527);
Av1EncoderTransformBlockState state = default;
Av1TransformBlockEncoder.EncodeIntraDcLossy(
workspace,
source8,
Stride,
reconstruction8,
Stride,
above8,
left8,
true,
true,
quantized,
transformSize,
Av1TransformType.DctDct,
73,
-1,
3,
Av1Plane.Y,
ref state);
Av1TransformBlockEncoder.EncodeIntraDcLossy(
workspace,
source10,
Stride,
reconstruction10,
Stride,
above10,
left10,
true,
true,
quantized,
transformSize,
Av1TransformType.DctDct,
73,
-1,
3,
Av1Plane.Y,
Av1BitDepth.TenBit,
ref state);
long before = GC.GetAllocatedBytesForCurrentThread();
for (int iteration = 0; iteration < 16; iteration++)
{
Av1TransformBlockEncoder.EncodeIntraDcLossy(
workspace,
source8,
Stride,
reconstruction8,
Stride,
above8,
left8,
true,
true,
quantized,
transformSize,
Av1TransformType.DctDct,
73,
-1,
3,
Av1Plane.Y,
ref state);
Av1TransformBlockEncoder.EncodeIntraDcLossy(
workspace,
source10,
Stride,
reconstruction10,
Stride,
above10,
left10,
true,
true,
quantized,
transformSize,
Av1TransformType.DctDct,
73,
-1,
3,
Av1Plane.Y,
Av1BitDepth.TenBit,
ref state);
}
Assert.Equal(0, GC.GetAllocatedBytesForCurrentThread() - before);
}
/// <summary>
/// Verifies that repeated maximum-transform block encoding uses only caller-owned workspaces.
/// </summary>
@ -179,6 +514,28 @@ public class Av1TransformBlockEncoderTests
}
}
private static void FillSource(Span<byte> source, int stride, int width, int height, int sampleMaximum)
{
for (int y = 0; y < height; y++)
{
for (int x = 0; x < width; x++)
{
source[(y * stride) + x] = (byte)(((y * 43) + (x * 71) + 29) % (sampleMaximum + 1));
}
}
}
private static void FillSource(Span<ushort> source, int stride, int width, int height, int sampleMaximum)
{
for (int y = 0; y < height; y++)
{
for (int x = 0; x < width; x++)
{
source[(y * stride) + x] = (ushort)(((y * 181) + (x * 313) + 97) % (sampleMaximum + 1));
}
}
}
private static void AssertEqual(ReadOnlySpan<int> expected, ReadOnlySpan<int> actual, int count)
{
for (int i = 0; i < count; i++)

Loading…
Cancel
Save