Browse Source

Add AV1 intra-block-copy mesh search

pull/2633/head
James Jackson-South 1 month ago
parent
commit
001c03bd8c
  1. 2
      HEIF_IMPLEMENTATION_PLAN.md
  2. 159
      src/ImageSharp/Formats/Heif/Av1/Motion/Av1IntraBlockCopySearchIndex.cs
  3. 136
      src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.Operator.cs
  4. 107
      tests/ImageSharp.Tests/Formats/Heif/Av1/Av1IntraBlockCopyTests.cs

2
HEIF_IMPLEMENTATION_PLAN.md

@ -859,7 +859,7 @@ Encoder verification contract:
- [~] Paired chroma palette clustering now preserves current libaom's squared two-component distance, first-centroid tie order, independently rounded U/V means, paired deterministic empty-cluster replacement, preceding-state retention on increased distortion, and 50-iteration limit. Keeping the source planes separate avoids interleave/deinterleave copies and improves on libaom's AVX2 ceiling with Vector512, Vector256, Vector128, then scalar dispatch through ImageSharp's shared vector-count helpers. Three independent tests cover exact paired convergence, midpoint initialization, 12-bit distance and index parity, untouched destination bounds, and every intrinsic tier. The exact Release test-project build reports 1,992 baseline warnings and zero errors; the focused three-case set, complete 8,934-case AVIF set, and complete 230-case HEIF set pass direct foreground net11 Release VSTest. Roslynk reports zero compiler errors and no touched-file analyzer warnings. Candidate integration and production activation remain in the open chroma-palette checkpoint.
- [~] Live paired chroma palette selection now follows current libaom's complete 2-through-8 color-size search, U-plane neighbor-cache snapping, stable U-ordered color pairs, shared U/V index map, implicit DCT-DCT transform, and strict rate-distortion winner replacement. It improves on speed-configured libaom by applying no early header-cost pruning, keeps planar U/V source data separate, and reuses the SIMD-first prediction, residual, transform, quantization, and reconstruction operators without allocator-backed candidate storage. The production tile regression proves both palette-mode probability branches, exact paired colors and indices, coefficient-free reconstruction, and nonempty syntax. The complete 58-case intra-superblock set, 8,935-case AVIF set, and 230-case HEIF set pass direct foreground net11 Release VSTest. The exact Release test-project build reports 1,992 baseline warnings and zero errors; Roslynk reports zero compiler errors and no touched-file analyzer warnings. Production frame activation remains the next checkpoint.
- [~] Production palette activation now matches current libaom's default good-quality screen detector: it scans only complete 16x16 luma blocks, normalizes high-bit-depth samples to eight bits, admits 2-through-4-color blocks, and uses the reference's strict greater-than-ten-percent frame-area threshold. A 256-bit stack bitset and a fifth-color early exit replace libaom's larger per-block histogram without changing the decision, allocation, or source precision. The adaptive sequence flag remains enabled, the frame flag is set before picture-state allocation, and intra-block copy remains disabled. Focused regressions prove strict-threshold equality, high-bit-depth normalization, five-color rejection, emitted frame-header activation, production decode, and generated payload retention. The exact Release test-project build reports 1,992 baseline warnings and zero errors; all 8,935 AVIF cases and all 230 HEIF cases pass direct foreground net11 Release VSTest. Current-main `aomdec` at `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727` accepts all 30 regenerated production payloads, including the 54-byte palette case. Roslynk reports zero compiler errors and no touched-file analyzer warnings.
- [~] Intra-block-copy rate accounting now uses the live frame-local flag and displacement-vector distributions without copying or adapting either context during candidate measurement. Displacement-vector costing and writing share one closed symbol operation over the exact current-libaom joint, sign, magnitude-class, class-zero, and integer-offset syntax; final mode evaluation applies libaom's 120/128 displacement-rate weight with nearest-integer rounding. Independent fixed costs cover all four joint states, both signs, class zero, and large offset classes before adaptive writes, followed by an encoder/decoder round trip through the same sequence. Encoder and decoder reference-vector derivation now share the exact eight-candidate spatial scan, independent nearest and outer-region ranking, top-right partition geometry, clamping, and tile-relative fallback. Selected vectors use a naturally aligned pair of signed 16-bit components packed into the existing picture-state owner only when intra-block copy is permitted; a 3840x2160 frame retains 130,560 vectors in 510 KiB while leaving the compact 8-byte mode allocation unchanged. The tile writer derives the same reference and emits the retained vector without another allocation or copy. Coefficient costing and writing now select the inter transform sets and frame-local probability tables required by intra-block copy; independent tests verify every legal symbol against the exact default inter distribution and round-trip full and reduced sets from 4x4 through 32x32. Legal 8x8 hash discovery now indexes every visible source origin, including unaligned origins, in libaom's coarse-to-fine insertion order with the same 256-candidate bucket cap. A separable rolling hash fills one packed picture-lifetime workspace before reconstruction, then reuses that workspace for integer candidate links; exact wide or SIMD block comparison rejects hash collisions, and SIMD variance uses libaom's eight-bit normalization at 8, 10, and 12 bits. Power-of-two bucket arrays scale down with small images and stop at the reference's 16-bit limit, avoiding libaom's fixed six-size pointer table; the 3840x2160 search index occupies about 32.2 MiB and introduces no additional owner or frame copy. Above and left search rectangles, integer displacement legality, strict tie order, and live raw displacement rate follow current libaom. Motion-candidate ranking uses libaom's undiscounted probability cost and exact variance-domain error-per-bit scaling, separately from the later 120/128 final-mode discount. The allocation-free full-pixel core now follows current libaom's NSTEP search: it clamps the spatial reference to each legal region, traverses the fixed 15-stage radii and site order, skips equivalent centered 210-pixel stages, repeats progressively shorter paths, and compares their winners in the normalized variance domain. Byte and high-bit-depth operators compute each 8x8 absolute difference with Vector128 before scalar fallback; high-bit-depth SAD remains in its native sample scale while its quantizer-derived rate multiplier uses libaom's normalized AC step. The exact net11 Release build reports 1,992 baseline warnings and zero errors; all 1,969 focused entropy and intra-block-copy cases, all 9,068 AV1 and AVIF cases, and all 189 non-AV1 HEIF cases pass through direct foreground VSTest, and Roslynk reports zero compiler errors with no touched-file analyzer diagnostics. The exhaustive mesh fallback, joint luma/chroma rate-distortion selection, production activation, and adaptive frame-flag clearing remain before intra-block copy can be enabled.
- [~] Intra-block-copy rate accounting now uses the live frame-local flag and displacement-vector distributions without copying or adapting either context during candidate measurement. Displacement-vector costing and writing share one closed symbol operation over the exact current-libaom joint, sign, magnitude-class, class-zero, and integer-offset syntax; final mode evaluation applies libaom's 120/128 displacement-rate weight with nearest-integer rounding. Independent fixed costs cover all four joint states, both signs, class zero, and large offset classes before adaptive writes, followed by an encoder/decoder round trip through the same sequence. Encoder and decoder reference-vector derivation now share the exact eight-candidate spatial scan, independent nearest and outer-region ranking, top-right partition geometry, clamping, and tile-relative fallback. Selected vectors use a naturally aligned pair of signed 16-bit components packed into the existing picture-state owner only when intra-block copy is permitted; a 3840x2160 frame retains 130,560 vectors in 510 KiB while leaving the compact 8-byte mode allocation unchanged. The tile writer derives the same reference and emits the retained vector without another allocation or copy. Coefficient costing and writing now select the inter transform sets and frame-local probability tables required by intra-block copy; independent tests verify every legal symbol against the exact default inter distribution and round-trip full and reduced sets from 4x4 through 32x32. Legal 8x8 hash discovery now indexes every visible source origin, including unaligned origins, in libaom's coarse-to-fine insertion order with the same 256-candidate bucket cap. A separable rolling hash fills one packed picture-lifetime workspace before reconstruction, then reuses that workspace for integer candidate links; exact wide or SIMD block comparison rejects hash collisions, and SIMD variance uses libaom's eight-bit normalization at 8, 10, and 12 bits. Power-of-two bucket arrays scale down with small images and stop at the reference's 16-bit limit, avoiding libaom's fixed six-size pointer table; the 3840x2160 search index occupies about 32.2 MiB and introduces no additional owner or frame copy. Above and left search rectangles, integer displacement legality, strict tie order, and live raw displacement rate follow current libaom. Motion-candidate ranking uses libaom's undiscounted probability cost and exact variance-domain error-per-bit scaling, separately from the later 120/128 final-mode discount. The allocation-free full-pixel core now follows current libaom's NSTEP search: it clamps the spatial reference to each legal region, traverses the fixed 15-stage radii and site order, skips equivalent centered 210-pixel stages, repeats progressively shorter paths, and compares their winners in the normalized variance domain. Paths above the speed-zero screen-content threshold continue through libaom's 256-pixel, one-pixel-step exhaustive mesh. Four adjacent byte or high-bit-depth candidates share each SIMD source load, strict row-major tie ordering is retained, and the final legal tail column remains searchable where libaom's current four-wide remainder loop omits it. Byte and high-bit-depth operators compute each 8x8 absolute difference with Vector128 before scalar fallback; high-bit-depth SAD remains in its native sample scale while its quantizer-derived rate multiplier uses libaom's normalized AC step. The exact net11 Release build reports 1,992 baseline warnings and zero errors; all 1,970 focused entropy and intra-block-copy cases, all 9,238 non-HEVC HEIF, AV1, and AVIF cases pass through direct foreground VSTest with tiered compilation disabled so runtime promotion bookkeeping cannot enter exact allocation-counter windows, and Roslynk reports zero compiler errors with no touched-file analyzer diagnostics. Joint luma/chroma rate-distortion selection, production activation, and adaptive frame-flag clearing remain before intra-block copy can be enabled.
- [x] The expanded checkpoint exposed a pre-existing transform-block test that asserted uninitialized pooled padding was zero. The test now initializes the complete physical luma plane with a sentinel and proves the block operation leaves both adjacent padding samples unchanged. The exact net11 Release rebuild remains at 1,005 baseline warnings and zero errors, the focused allocator-order set passes 30 of 30 cases, and the complete HEIF/AV1 namespace passes 8,859 of 8,859 direct VSTest cases with zero failures or skips.
- [x] Combined-frame OBU output now counts the byte-aligned frame and tile-group headers, non-final tile-size fields, and owned tile payloads before emitting the OBU size. It retains only the small allocator-owned header scratch and writes each entropy-coded tile span directly from its detached owner, removing the second file-sized allocator rent and complete-payload copy. A 64 KiB regression proves exactly one sub-payload-sized byte rent with a balanced return and verifies the exact streamed tile tail; the existing two-tile round trip proves size-prefix and ordering parity. The focused writer and production-frame set passes 32 of 32 direct net11 VSTest cases, current-main `aomdec` accepts all 29 generated native-format payloads, and the complete HEIF/AV1 namespace passes 8,860 of 8,860 cases with zero failures or skips.
- [x] Finalized fixed-block decisions now set the block-level transform-skip flag only when every retained luma and coded chroma transform has zero EOB, matching current libaom's conjunction of per-plane skip state. The previous always-false flag produced legal but redundant non-skip and zero-coefficient syntax. Monochrome and 4:2:0 regressions prove both branches from actual coefficient state; the focused decision and production-frame set passes 32 of 32 direct net11 VSTest cases. Current-main `aomdec` accepts all 29 regenerated payloads, the recorded decoded-frame MD5s are unchanged, and affected 16x16 constant 8-bit and 10-bit payloads are one byte smaller. The complete HEIF/AV1 namespace passes 8,862 of 8,862 cases with zero failures or skips.

159
src/ImageSharp/Formats/Heif/Av1/Motion/Av1IntraBlockCopySearchIndex.cs

@ -20,6 +20,9 @@ internal readonly struct Av1IntraBlockCopySearchIndex
private const int MaximumFullPixelSearchOffset = (1 << 10) - 1;
private const int MinimumFullPixelMotionVector = -(1 << 11) + 1;
private const int MaximumFullPixelMotionVector = (1 << 11) - 1;
private const int ExhaustiveSearchRange = 256;
private const int ExhaustiveSearchThreshold = 1 << 12;
private const int ExhaustiveSearchBatchSize = 4;
private const uint HorizontalHashMultiplier = 257;
private const uint VerticalHashMultiplier = 65599;
private static readonly uint HorizontalLeadingWeight = GetLeadingWeight(HorizontalHashMultiplier);
@ -89,6 +92,21 @@ internal readonly struct Av1IntraBlockCopySearchIndex
Buffer2DRegion<TSample> reconstruction,
Point predictionOrigin);
/// <summary>
/// Gets four sums of absolute differences for horizontally adjacent reconstructed predictors.
/// </summary>
/// <param name="source">The coded source plane.</param>
/// <param name="sourceOrigin">The source block origin.</param>
/// <param name="reconstruction">The reconstructed luma plane.</param>
/// <param name="firstPredictionOrigin">The first of four horizontally adjacent predictor origins.</param>
/// <param name="sums">Storage receiving the four unnormalized absolute differences.</param>
public static abstract void GetFourSumsOfAbsoluteDifferences(
Buffer2DRegion<TSample> source,
Point sourceOrigin,
Buffer2DRegion<TSample> reconstruction,
Point firstPredictionOrigin,
Span<int> sums);
/// <summary>
/// Gets the normalized 8x8 variance between a source block and reconstructed predictor.
/// </summary>
@ -734,6 +752,39 @@ internal readonly struct Av1IntraBlockCopySearchIndex
shortenedBy += additionalCenterSteps;
}
// Intra-block copy is a screen-content tool. Scaling the encoder's 1 << 20 full-search threshold
// by the 8x8 block area yields this normalized variance-domain trigger.
if (bestCost > ExhaustiveSearchThreshold)
{
Point candidate = SearchExhaustiveMesh<TSample, TOperation>(
source,
reconstruction,
blockOrigin,
writer,
reference,
sadPerBit,
minimumColumnOffset,
minimumRowOffset,
maximumColumnOffset,
maximumRowOffset,
best);
int candidateCost = GetVarianceCost<TSample, TOperation>(
source,
reconstruction,
blockOrigin,
writer,
reference,
bitDepth,
rateMultiplier,
candidate);
if (candidateCost < bestCost)
{
best = candidate;
}
}
return best;
}
@ -833,6 +884,114 @@ internal readonly struct Av1IntraBlockCopySearchIndex
}
}
private static Point SearchExhaustiveMesh<TSample, TOperation>(
Buffer2DRegion<TSample> source,
Buffer2DRegion<TSample> reconstruction,
Point blockOrigin,
Av1SymbolEncoder writer,
Av1MotionVector reference,
int sadPerBit,
int minimumColumnOffset,
int minimumRowOffset,
int maximumColumnOffset,
int maximumRowOffset,
Point start)
where TSample : unmanaged
where TOperation : struct, ISearchOperation<TSample>
{
Point best = start;
int bestCost = GetSadCost<TSample, TOperation>(
source,
reconstruction,
blockOrigin,
writer,
reference,
sadPerBit,
start);
int startColumn = Math.Max(-ExhaustiveSearchRange, minimumColumnOffset - start.X);
int endColumn = Math.Min(ExhaustiveSearchRange, maximumColumnOffset - start.X);
int startRow = Math.Max(-ExhaustiveSearchRange, minimumRowOffset - start.Y);
int endRow = Math.Min(ExhaustiveSearchRange, maximumRowOffset - start.Y);
Span<int> sumsOfAbsoluteDifferences = stackalloc int[ExhaustiveSearchBatchSize];
for (int row = startRow; row <= endRow; row++)
{
int column = startColumn;
for (; column <= endColumn - (ExhaustiveSearchBatchSize - 1); column += ExhaustiveSearchBatchSize)
{
Point firstCandidate = new(start.X + column, start.Y + row);
Point firstPredictionOrigin = new(
blockOrigin.X + firstCandidate.X,
blockOrigin.Y + firstCandidate.Y);
// Four adjacent candidates share the source load and row traversal, matching the batch width
// used by the native full-resolution pass without allocating temporary candidate buffers.
TOperation.GetFourSumsOfAbsoluteDifferences(
source,
blockOrigin,
reconstruction,
firstPredictionOrigin,
sumsOfAbsoluteDifferences);
for (int i = 0; i < ExhaustiveSearchBatchSize; i++)
{
int sumOfAbsoluteDifferences = sumsOfAbsoluteDifferences[i];
if (sumOfAbsoluteDifferences >= bestCost)
{
continue;
}
Point candidate = new(firstCandidate.X + i, firstCandidate.Y);
Av1MotionVector vector = new(candidate.Y * 8, candidate.X * 8);
int rate = writer.GetDisplacementVectorSearchCost(vector, reference);
int candidateCost = Av1RateDistortion.GetMotionSearchSadCost(
sadPerBit,
rate,
sumOfAbsoluteDifferences);
// Strict replacement preserves the first row-major candidate when costs tie.
if (candidateCost < bestCost)
{
bestCost = candidateCost;
best = candidate;
}
}
}
// The SIMD batch width is only a traversal optimization; every legal tail column remains searchable.
for (; column <= endColumn; column++)
{
Point candidate = new(start.X + column, start.Y + row);
Point predictionOrigin = new(blockOrigin.X + candidate.X, blockOrigin.Y + candidate.Y);
int sumOfAbsoluteDifferences = TOperation.GetSumOfAbsoluteDifferences(
source,
blockOrigin,
reconstruction,
predictionOrigin);
if (sumOfAbsoluteDifferences >= bestCost)
{
continue;
}
Av1MotionVector vector = new(candidate.Y * 8, candidate.X * 8);
int rate = writer.GetDisplacementVectorSearchCost(vector, reference);
int candidateCost = Av1RateDistortion.GetMotionSearchSadCost(
sadPerBit,
rate,
sumOfAbsoluteDifferences);
if (candidateCost < bestCost)
{
bestCost = candidateCost;
best = candidate;
}
}
}
return best;
}
private static int GetSadCost<TSample, TOperation>(
Buffer2DRegion<TSample> source,
Buffer2DRegion<TSample> reconstruction,

136
src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.Operator.cs

@ -382,6 +382,74 @@ internal static partial class Av1IntraSuperblockEncoder
return sum;
}
/// <inheritdoc/>
public static void GetFourSumsOfAbsoluteDifferences(
Buffer2DRegion<byte> source,
Point sourceOrigin,
Buffer2DRegion<byte> reconstruction,
Point firstPredictionOrigin,
Span<int> sums)
{
int sum0 = 0;
int sum1 = 0;
int sum2 = 0;
int sum3 = 0;
if (Vector128.IsHardwareAccelerated)
{
for (int row = 0; row < 8; row++)
{
ReadOnlySpan<byte> sourceRow = source.DangerousGetRowSpan(sourceOrigin.Y + row)[sourceOrigin.X..];
ReadOnlySpan<byte> predictionRow =
reconstruction.DangerousGetRowSpan(firstPredictionOrigin.Y + row)[firstPredictionOrigin.X..];
// The four candidates reuse one widened source vector; only their overlapping predictor
// windows are loaded separately before the packed absolute-difference reductions.
Vector128<short> sourceSamples =
Vector128.WidenLower(Vector128.CreateScalarUnsafe(MemoryMarshal.Read<ulong>(sourceRow)).AsByte()).AsInt16();
Vector128<short> prediction0 =
Vector128.WidenLower(Vector128.CreateScalarUnsafe(MemoryMarshal.Read<ulong>(predictionRow)).AsByte()).AsInt16();
Vector128<short> prediction1 =
Vector128.WidenLower(Vector128.CreateScalarUnsafe(MemoryMarshal.Read<ulong>(predictionRow[1..])).AsByte()).AsInt16();
Vector128<short> prediction2 =
Vector128.WidenLower(Vector128.CreateScalarUnsafe(MemoryMarshal.Read<ulong>(predictionRow[2..])).AsByte()).AsInt16();
Vector128<short> prediction3 =
Vector128.WidenLower(Vector128.CreateScalarUnsafe(MemoryMarshal.Read<ulong>(predictionRow[3..])).AsByte()).AsInt16();
sum0 += Vector128.Sum(Vector128.Abs(sourceSamples - prediction0));
sum1 += Vector128.Sum(Vector128.Abs(sourceSamples - prediction1));
sum2 += Vector128.Sum(Vector128.Abs(sourceSamples - prediction2));
sum3 += Vector128.Sum(Vector128.Abs(sourceSamples - prediction3));
}
}
else
{
for (int row = 0; row < 8; row++)
{
ReadOnlySpan<byte> sourceRow = source.DangerousGetRowSpan(sourceOrigin.Y + row)[sourceOrigin.X..];
ReadOnlySpan<byte> predictionRow =
reconstruction.DangerousGetRowSpan(firstPredictionOrigin.Y + row)[firstPredictionOrigin.X..];
for (int column = 0; column < 8; column++)
{
int sourceSample = sourceRow[column];
sum0 += Math.Abs(sourceSample - predictionRow[column]);
sum1 += Math.Abs(sourceSample - predictionRow[column + 1]);
sum2 += Math.Abs(sourceSample - predictionRow[column + 2]);
sum3 += Math.Abs(sourceSample - predictionRow[column + 3]);
}
}
}
sums[0] = sum0;
sums[1] = sum1;
sums[2] = sum2;
sums[3] = sum3;
}
/// <inheritdoc/>
public static int GetVariance(
Buffer2DRegion<byte> source,
@ -798,6 +866,74 @@ internal static partial class Av1IntraSuperblockEncoder
return sum;
}
/// <inheritdoc/>
public static void GetFourSumsOfAbsoluteDifferences(
Buffer2DRegion<ushort> source,
Point sourceOrigin,
Buffer2DRegion<ushort> reconstruction,
Point firstPredictionOrigin,
Span<int> sums)
{
int sum0 = 0;
int sum1 = 0;
int sum2 = 0;
int sum3 = 0;
if (Vector128.IsHardwareAccelerated)
{
for (int row = 0; row < 8; row++)
{
ReadOnlySpan<ushort> sourceRow = source.DangerousGetRowSpan(sourceOrigin.Y + row)[sourceOrigin.X..];
ReadOnlySpan<ushort> predictionRow =
reconstruction.DangerousGetRowSpan(firstPredictionOrigin.Y + row)[firstPredictionOrigin.X..];
// Signed 16-bit lanes preserve every AV1 sample difference while four horizontally adjacent
// candidates reuse the same source load and stay packed through horizontal reduction.
Vector128<short> sourceSamples =
Vector128.LoadUnsafe(ref MemoryMarshal.GetReference(sourceRow)).AsInt16();
Vector128<short> prediction0 =
Vector128.LoadUnsafe(ref MemoryMarshal.GetReference(predictionRow)).AsInt16();
Vector128<short> prediction1 =
Vector128.LoadUnsafe(ref MemoryMarshal.GetReference(predictionRow[1..])).AsInt16();
Vector128<short> prediction2 =
Vector128.LoadUnsafe(ref MemoryMarshal.GetReference(predictionRow[2..])).AsInt16();
Vector128<short> prediction3 =
Vector128.LoadUnsafe(ref MemoryMarshal.GetReference(predictionRow[3..])).AsInt16();
sum0 += Vector128.Sum(Vector128.Abs(sourceSamples - prediction0));
sum1 += Vector128.Sum(Vector128.Abs(sourceSamples - prediction1));
sum2 += Vector128.Sum(Vector128.Abs(sourceSamples - prediction2));
sum3 += Vector128.Sum(Vector128.Abs(sourceSamples - prediction3));
}
}
else
{
for (int row = 0; row < 8; row++)
{
ReadOnlySpan<ushort> sourceRow = source.DangerousGetRowSpan(sourceOrigin.Y + row)[sourceOrigin.X..];
ReadOnlySpan<ushort> predictionRow =
reconstruction.DangerousGetRowSpan(firstPredictionOrigin.Y + row)[firstPredictionOrigin.X..];
for (int column = 0; column < 8; column++)
{
int sourceSample = sourceRow[column];
sum0 += Math.Abs(sourceSample - predictionRow[column]);
sum1 += Math.Abs(sourceSample - predictionRow[column + 1]);
sum2 += Math.Abs(sourceSample - predictionRow[column + 2]);
sum3 += Math.Abs(sourceSample - predictionRow[column + 3]);
}
}
}
sums[0] = sum0;
sums[1] = sum1;
sums[2] = sum2;
sums[3] = sum3;
}
/// <inheritdoc/>
public static int GetVariance(
Buffer2DRegion<ushort> source,

107
tests/ImageSharp.Tests/Formats/Heif/Av1/Av1IntraBlockCopyTests.cs

@ -413,15 +413,99 @@ public class Av1IntraBlockCopyTests
Assert.Equal(120, candidates[0].Column);
}
/// <summary>
/// Verifies that the exhaustive mesh recovers an exact match outside every centered NSTEP search site.
/// </summary>
[Fact]
public void PixelSearchFallsBackToExhaustiveMesh()
{
const int Width = 640;
const int Height = 256;
const int QIndex = 23;
Point blockOrigin = new(0, 128);
Point predictionOrigin = new(256, 8);
ObuSequenceHeader sequenceHeader = CreateSequenceHeader();
ObuFrameHeader frameHeader = CreateFrameHeader();
frameHeader.AllowScreenContentTools = true;
frameHeader.AllowIntraBlockCopy = true;
using Av1EncoderPictureBuffer pictureBuffer = new(
Configuration.Default,
sequenceHeader,
frameHeader,
Width,
Height);
using Av1EncoderFrameBuffer<byte> source = new(
Configuration.Default,
Width,
Height,
8,
Av1ColorFormat.Yuv400,
0,
0);
using Av1EncoderFrameBuffer<byte> reconstruction = new(
Configuration.Default,
Width,
Height,
8,
Av1ColorFormat.Yuv400,
0,
0);
Buffer2DRegion<byte> sourceLuma = source.Frame.View.GetPlane(Av1Plane.Y);
Buffer2DRegion<byte> reconstructionLuma = reconstruction.Frame.View.GetPlane(Av1Plane.Y);
for (int row = 0; row < Height; row++)
{
sourceLuma.DangerousGetRowSpan(row).Clear();
reconstructionLuma.DangerousGetRowSpan(row).Clear();
}
for (int row = 0; row < 8; row++)
{
Span<byte> sourceRow = sourceLuma.DangerousGetRowSpan(blockOrigin.Y + row).Slice(blockOrigin.X, 8);
Span<byte> predictionRow = reconstructionLuma.DangerousGetRowSpan(predictionOrigin.Y + row).Slice(predictionOrigin.X, 8);
for (int column = 0; column < 8; column++)
{
byte value = (byte)((row + column) % 2 == 0 ? 255 : 0);
sourceRow[column] = value;
predictionRow[column] = value;
}
}
Buffer2DRegion<byte> codedSourceLuma = source.Frame.CodedView.GetPlane(Av1Plane.Y);
Buffer2DRegion<byte> codedReconstructionLuma = reconstruction.Frame.CodedView.GetPlane(Av1Plane.Y);
using Av1SymbolEncoder writer = new(Configuration.Default, 64, QIndex);
Span<Av1MotionVector> candidates = stackalloc Av1MotionVector[2];
int candidateCount = pictureBuffer.Picture.IntraBlockCopySearch
.FindPixelCandidates<byte, Av1IntraSuperblockEncoder.ByteOperator>(
codedSourceLuma,
codedReconstructionLuma,
blockOrigin,
new Av1TileInfo(0, 0, frameHeader),
sequenceHeader,
writer,
new Av1MotionVector(-64, 0),
QIndex,
Av1RateDistortion.GetKeyFrameRateMultiplier(QIndex, Av1BitDepth.EightBit),
candidates);
Assert.Equal(1, candidateCount);
Assert.Equal(-960, candidates[0].Row);
Assert.Equal(2048, candidates[0].Column);
}
/// <summary>
/// Verifies high-bit-depth SIMD variance normalization against the eight-bit search domain.
/// </summary>
[Fact]
public void SearchVarianceMatchesTwelveBitReference()
{
const int Width = 11;
using Av1EncoderFrameBuffer<ushort> source = new(
Configuration.Default,
8,
Width,
8,
12,
Av1ColorFormat.Yuv400,
@ -430,7 +514,7 @@ public class Av1IntraBlockCopyTests
using Av1EncoderFrameBuffer<ushort> reconstruction = new(
Configuration.Default,
8,
Width,
8,
12,
Av1ColorFormat.Yuv400,
@ -463,6 +547,25 @@ public class Av1IntraBlockCopyTests
Point.Empty,
Av1BitDepth.TwelveBit);
Span<int> fourSumsOfAbsoluteDifferences = stackalloc int[4];
Av1IntraSuperblockEncoder.UInt16Operator.GetFourSumsOfAbsoluteDifferences(
sourceLuma,
Point.Empty,
reconstructionLuma,
Point.Empty,
fourSumsOfAbsoluteDifferences);
for (int candidate = 0; candidate < fourSumsOfAbsoluteDifferences.Length; candidate++)
{
int expected = Av1IntraSuperblockEncoder.UInt16Operator.GetSumOfAbsoluteDifferences(
sourceLuma,
Point.Empty,
reconstructionLuma,
new Point(candidate, 0));
Assert.Equal(expected, fourSumsOfAbsoluteDifferences[candidate]);
}
Assert.Equal(1600, sumOfAbsoluteDifferences);
Assert.Equal(16, variance);
}

Loading…
Cancel
Save