diff --git a/HEIF_IMPLEMENTATION_PLAN.md b/HEIF_IMPLEMENTATION_PLAN.md
index dc221f0cb..bef99942b 100644
--- a/HEIF_IMPLEMENTATION_PLAN.md
+++ b/HEIF_IMPLEMENTATION_PLAN.md
@@ -29,7 +29,7 @@ Checkboxes may be marked complete only when the implementation and the verificat
## Delivery dashboard
-Last reconciled with the source tree on 2026-08-29 against production checkpoint `beb2ab86ce5e4052c667df30af8a48a860814031`. Committed checkpoints include the AV1 transform architecture, OBU framing, intra-block copy, 12-profile reconstruction matrix, layered-item properties, layered reference/header/CDF/motion-field state, inter-frame intra blocks, SIMD-first translational prediction, complete single-reference inter reconstruction, compound reference trees and modes, paired reference-MV derivation, reference-dependent bounded sequence decoding, allocation-free SIMD-first equal averaging, selected inter-intra prediction, selectable compound blending, OBMC, scaled-reference reconstruction, local warped prediction, non-translational global prediction, official motion-vector conformance, selected spatial-layer presentation, and progressive color and auxiliary-alpha conformance. This dashboard is the authoritative delivery order. The detailed phase checklists below provide subsystem evidence; they do not override the current-stage marker or permit work to skip ahead.
+Last reconciled with the source tree on 2026-08-29 against production checkpoint `ef93d584511055f8e91e8d801662c44a5b0984d9`. Committed checkpoints include the AV1 transform architecture, OBU framing, intra-block copy, 12-profile reconstruction matrix, corrected twelve-bit inverse-transform SIMD arithmetic, layered-item properties, layered reference/header/CDF/motion-field state, inter-frame intra blocks, SIMD-first translational prediction, complete single-reference inter reconstruction, compound reference trees and modes, paired reference-MV derivation, reference-dependent bounded sequence decoding, allocation-free SIMD-first equal averaging, selected inter-intra prediction, selectable compound blending, OBMC, scaled-reference reconstruction, local warped prediction, non-translational global prediction, official motion-vector conformance, selected spatial-layer presentation, and progressive color and auxiliary-alpha conformance. This dashboard is the authoritative delivery order. The detailed phase checklists below provide subsystem evidence; they do not override the current-stage marker or permit work to skip ahead.
Commit `1c58d855f70b024170ced9eb0a7005f0f9c955ad` records the complete official motion-vector conformance checkpoint. The official IVF has SHA-1 `F064290D7FCD3B3DE19020E8AEC6C43C88D3A505`, matching the pinned libaom test-data manifest, and SHA-256 `222A9050059B254DAB17CFB802FF829C778E3F93AF18961A622C8268576C1395`; its pinned-libaom Y4M has SHA-256 `D97AC78C81782CF1507549368047769DC677DBE205706458D1EE9C807DE6EC78`. Both source targets build with zero warnings and errors, the `net10.0` test-project analyzer build completes with zero errors and 1,014 pre-existing repository warnings, Roslynk reports zero compiler errors, 3,985 focused `net10.0` cases pass without failures or skips, and `git diff --check` is clean.
@@ -37,7 +37,9 @@ Commit `019ac5648b4380a6ff4865072c96fde29ce09fad` records the complete selected-
Commit `beb2ab86ce5e4052c667df30af8a48a860814031` records the complete progressive color and auxiliary-alpha conformance checkpoint. Both pinned progressive AVIF fixtures pass their exact native-plane and pinned-libavif presentation comparisons through normal dispatch, AVX-512-disabled, AVX-disabled, and scalar `FeatureTestRunner` configurations. The YUV444-alpha fixture also compares every composed alpha sample with the independent final native auxiliary plane. Both complete production presentation paths pass under a 1,024-byte constrained tracked allocator with every allocation returned exactly once. Both source targets build with zero warnings and errors, the `net10.0` test-project analyzer build completes with zero errors and 1,014 pre-existing repository warnings, Roslynk reports zero compiler errors, 156 focused reconstruction and color cases pass without failures or skips, and `git diff --check` is clean.
-The working tree contains a complete, verified twelve-bit inverse-transform arithmetic checkpoint awaiting commit. ADST4 retains pinned libaom's signed 32-bit sine products and factorized sums and widens only its terminal rounding; Identity4 and Identity16 widen their fixed-point product and rounding bias only for the 20-bit twelve-bit row stage. The established 8/10-bit and twelve-bit column paths remain unchanged. Exact boundary vectors cover both 128-bit and 256-bit operators through normal, AVX-512-disabled, AVX-disabled, and scalar `FeatureTestRunner` configurations. Both source targets build with zero warnings and errors, the `net10.0` test-project analyzer build completes with zero errors and 1,014 pre-existing repository warnings, Roslynk reports zero compiler errors, 510 focused inverse-transform cases and 44 production reconstruction cases pass without failures or skips, and `git diff --check` is clean.
+Commit `ef93d584511055f8e91e8d801662c44a5b0984d9` records the complete twelve-bit inverse-transform arithmetic checkpoint. ADST4 retains pinned libaom's signed 32-bit sine products and factorized sums and widens only its terminal rounding; Identity4 and Identity16 widen their fixed-point product and rounding bias only for the 20-bit twelve-bit row stage. The established 8/10-bit and twelve-bit column paths remain unchanged. Exact boundary vectors cover both 128-bit and 256-bit operators through normal, AVX-512-disabled, AVX-disabled, and scalar `FeatureTestRunner` configurations. Both source targets build with zero warnings and errors, the `net10.0` test-project analyzer build completes with zero errors and 1,014 pre-existing repository warnings, Roslynk reports zero compiler errors, 510 focused inverse-transform cases and 44 production reconstruction cases pass without failures or skips, and `git diff --check` is clean.
+
+The working tree contains a complete, verified official syntax-coverage and predictor-architecture checkpoint awaiting commit. The official all-intra, CDF-update, and temporal motion-field IVF files have SHA-1 `A9F7EA6312A533CC6426A6145EDD190D45813C37`, `AFCA5502A489692B0A3C120370B0F43B8FC572A1`, and `B48A717C7C003B8DD23C3C2CAED1AC673380FDB3`, exactly matching the pinned libaom manifest. Their SHA-256 values are `5FCD265FD9F9BDD0D3179340B4C4532F1422CA5E5D97741C7481B84CB5DC122F`, `14A3DBF537B6BF15EFC003182D9916D61438C93624A8BD26E6E3AE7EAF33EA82`, and `B59BF9586D8546DFDA81DFEC4EE4E32CEB502C9D22412AB0B63A2ABB534A1F14`; the pinned-libaom Y4M references have SHA-256 `1211EBEFBC9CCEF9ED19BE4CCE3F807D69FFFE338E95CCA1B5F4CA8023482175`, `4FBFF73FF0DE2D9084DAE557D1D4BD677B0486516525BF4D327D2D795D5A7779`, and `F7DB607694818C19E62FD9A27F53E1A3E2D00B72C39C0430C1B26399CC76777D`. Exact native-plane comparison covers all 39 all-intra frames, every intra mode, seven selected transform types, both tile-local and frame-end adaptive CDF publication, all four temporal motion-field frames under normal/scalar dispatch, and constrained tracked motion-field allocation with exactly-once returns. Every distinct AV1 predictor traversal now owns a family-named JPEG-style static operator contract and SIMD traversal instead of nesting separate predictors beneath broad intra/inter families. Compound reference convolution, equal averaging, distance weighting, alpha-mask blending, difference-weighted mask construction, and each intermediate reconstruction or mask traversal have separate family owners and matching `.Operator.cs` contracts; inter-intra mask construction has its own non-operator builder. Both Release source targets build with zero warnings and errors, Roslynk reports zero compiler errors, 179 focused prediction, reconstruction, reference, ownership, and lifecycle cases pass without failures or skips, and `git diff --check` is clean.
Status meanings:
@@ -49,7 +51,7 @@ Status meanings:
Current development stage: **Stage 3 — complete AV1 still-image decoding.** The decoder retains reference/header/CDF/motion-field state, derives frame-level skip-mode references, consumes temporal segment prediction, decodes intra-coded blocks inside inter frames, and reconstructs translational single-reference, compound, inter-intra, OBMC, scaled-reference, local warped, and non-translational global prediction before residual traversal. Commit `1c58d855f70b024170ced9eb0a7005f0f9c955ad` adds exact official four-frame coverage through all ordinary inter modes, all three motion modes, every regular/smooth/sharp dual-filter pair, sub-8x8 chroma prediction, and no-round compound intermediates. Commit `019ac5648b4380a6ff4865072c96fde29ce09fad` adds exact selected-layer native reconstruction and scaled presentation. Neither AV1 nor HEVC production encoding is implemented.
-Immediate checkpoint: **commit the verified twelve-bit inverse-transform arithmetic correction before advancing.** The corrected operators match pinned libaom at the exact overflow boundaries, preserve narrower stage paths, and pass the complete focused transform and production reconstruction evidence.
+Immediate checkpoint: **inventory and remove the next remaining production-reachable unsupported valid AV1 still-image syntax path.** Trace every explicit rejection through the decoder, distinguish malformed or out-of-scope syntax from valid still-image behavior, select the first valid gap in source order, and close it with exact independent native-plane and presentation evidence before advancing.
| Order | Delivery stage | State | Delivered state | Gate that remains open |
| --- | --- | --- | --- | --- |
@@ -107,7 +109,8 @@ Immediate checkpoint: **commit the verified twelve-bit inverse-transform arithme
- [x] Return the explicitly selected spatial layer or the final displayed layer, keeping reference reconstruction separate from display-only film grain. The essential-`lsel` production fixture reconstructs the selected 40x40 YUV444 base exactly, scales native component planes to the 80x80 item extent with pinned-libyuv integer rounding, and matches pinned-libavif RGBA presentation under normal/scalar dispatch. The committed final-layer fixture and film-grain matrix remain exact. Constrained tracked allocation returns every short-lived presentation plane exactly once. Commit `019ac5648b4380a6ff4865072c96fde29ce09fad` records the checkpoint.
- [x] Verify color and auxiliary-alpha output exactly against both pinned libavif progressive fixtures under normal SIMD dispatch and all required `FeatureTestRunner` fallbacks. Both fixtures pass exact public presentation under normal, AVX-512-disabled, AVX-disabled, and scalar dispatch; the alpha-bearing fixture additionally matches every composed alpha sample with its pinned native auxiliary plane. Both full production paths pass under a 1,024-byte constrained tracked allocator with exactly-once returns. Both Release source targets build with zero warnings and errors, the test-project analyzer build completes with zero errors and 1,014 pre-existing warnings, Roslynk reports zero compiler errors, and all 156 focused reconstruction and color cases pass without failures or skips.
- [x] Correct the audited 12-bit inverse ADST4, Identity4, and Identity16 SIMD arithmetic by widening only the libaom-widened multiply/accumulate operations, with exact conformant-range vectors and `FeatureTestRunner` coverage. Both 128-bit and 256-bit operators match pinned outputs through every required hardware fallback; 510 focused inverse-transform cases, 44 production reconstruction cases, both zero-warning source builds, the zero-error analyzer build, zero Roslynk compiler errors, and clean `git diff --check` complete the verification.
-- [ ] **Queued until the twelve-bit inverse-transform checkpoint commit:** continue inventorying and removing every remaining valid AV1 still-image unsupported branch, adding exact independent compression-tool fixtures to the profile-matrix regression gate.
+- [x] Verify the official libaom all-intra, CDF-update, and temporal motion-field sequences through exact native output, selected syntax coverage, normal/scalar dispatch, constrained tracked allocation, and the family-owned JPEG-style predictor operator architecture.
+- [ ] **Current:** continue inventorying and removing every remaining valid AV1 still-image unsupported branch, adding exact independent compression-tool fixtures to the profile-matrix regression gate.
- [ ] Complete the remaining HEVC still-image profile and Range Extensions matrix with exact independent native-plane and presentation evidence.
- [ ] Close shared decoded presentation, ICC, alpha, grid, transform, metadata, and animated AV1/HEVC decode gates.
- [ ] Implement and independently verify real AV1/AVIF still encoding.
@@ -819,6 +822,7 @@ No valid HEVC or AV1 color, compression, or bit-depth row may remain `unsupporte
## Working rules for implementation
- Keep changes vertical and reviewable. A slice should add one behavior, its focused tests, independent evidence, and any required notice update.
+- Keep every AV1 prediction family on the established JPEG color-converter operator architecture. Each distinct traversal contract owns a family-named predictor type; its `.Operator.cs` defines the static interface, and its family-named files own the closed generic widest-to-narrowest SIMD traversal. Semantic `readonly struct` operators implement scalar, `Vector128`, `Vector256`, and `Vector512` arithmetic through that contract. Only modes which share the same traversal and contract may share a predictor family; do not nest a separate predictor beneath a broad intra/inter family or create hardware-width-specific class hierarchies.
- Design SIMD-suitable codec work SIMD-first. Establish vector-friendly storage, operator boundaries, scratch ownership, traversal, every applicable lane width, and benchmark-gated dispatch before implementing the equivalent scalar fallback; never build a scalar production architecture and bolt SIMD onto it later.
- Inspect every owning method and upstream invariant before adding guards. Validate external file data at the parser/model boundary and rely on those established invariants internally.
- Do not extract one-use helpers merely to label code. Extract shared primitives only when they have genuine reuse or remove substantial complexity.
@@ -841,7 +845,7 @@ The dashboard and immediate execution queue define the remaining critical path.
- [x] Implement and independently verify non-translational global prediction with a genuine traced bounded sequence, exact native and presentation comparisons, constrained allocation, normal/scalar dispatch, and direct 8/10/12-bit compound production coverage. Commit `c5637ea0187df35b385bf43e2fe85cd955f01099` records the checkpoint.
- [x] Complete progressive color and auxiliary-alpha verification through exact native and presentation comparisons, every required `FeatureTestRunner` fallback, and constrained tracked allocation.
- [x] Correct the audited 12-bit inverse ADST4, Identity4, and Identity16 arithmetic through exact pinned boundary vectors and SIMD/scalar feature isolation.
-- [ ] **Queued until the twelve-bit inverse-transform checkpoint commit:** remove every other unsupported valid AV1 still-image syntax path and prove the complete AVIF decode matrix with independent inputs and scalar/SIMD parity.
+- [ ] **Current:** remove every other unsupported valid AV1 still-image syntax path and prove the complete AVIF decode matrix with independent inputs and scalar/SIMD parity.
- [ ] Close Phase 4 by completing the remaining HEVC profile and Range Extensions matrix with exact native-plane and presented-image evidence.
- [ ] Close Phase 5 and the decode portion of the bounded sequence ledger: color, ICC, alpha, grids, presentation transforms, reference-dependent samples, and complete animated AVIF/HEIC decode.
- [ ] Close the still-image portions of Phases 0, 1, and 2 that remain as release gates: documentation, provenance, public format boundaries, API review, parser hardening, and malformed-input coverage.
diff --git a/shared-infrastructure b/shared-infrastructure
index 03471c6b4..a835a9d74 160000
--- a/shared-infrastructure
+++ b/shared-infrastructure
@@ -1 +1 @@
-Subproject commit 03471c6b458a2c11a0b1df1f5bf2777c8c773b99
+Subproject commit a835a9d74e82b2d32b580a7902eb2699ebc47098
diff --git a/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1DcIntraPredictor.Operator.cs b/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1DcIntraPredictor.Operator.cs
new file mode 100644
index 000000000..8dedd0fc4
--- /dev/null
+++ b/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1DcIntraPredictor.Operator.cs
@@ -0,0 +1,434 @@
+// Copyright (c) Six Labors.
+// Licensed under the Six Labors Split License.
+
+using System.Runtime.CompilerServices;
+using System.Runtime.InteropServices;
+using System.Runtime.Intrinsics;
+
+namespace SixLabors.ImageSharp.Formats.Heif.Av1.Prediction;
+
+///
+/// Defines reference reduction and rounded mean arithmetic for AV1 DC intra prediction.
+///
+internal static class Av1DcIntraPredictor
+{
+ ///
+ /// Defines scalar and SIMD reference reduction for AV1 DC intra prediction.
+ ///
+ internal interface IDcPredictionOperator
+ {
+ ///
+ /// Sums one 8-bit reference sample.
+ ///
+ /// The reference sample.
+ /// The sample value.
+ public static abstract int Sum(byte sample);
+
+ ///
+ /// Sums sixteen 8-bit reference samples.
+ ///
+ /// The reference samples.
+ /// The exact sum.
+ public static abstract int Sum(Vector128 samples);
+
+ ///
+ /// Sums thirty-two 8-bit reference samples.
+ ///
+ /// The reference samples.
+ /// The exact sum.
+ public static abstract int Sum(Vector256 samples);
+
+ ///
+ /// Sums sixty-four 8-bit reference samples.
+ ///
+ /// The reference samples.
+ /// The exact sum.
+ public static abstract int Sum(Vector512 samples);
+
+ ///
+ /// Sums one high-bit-depth reference sample.
+ ///
+ /// The reference sample.
+ /// The sample value.
+ public static abstract int Sum(short sample);
+
+ ///
+ /// Sums eight high-bit-depth reference samples.
+ ///
+ /// The reference samples.
+ /// The exact sum.
+ public static abstract int Sum(Vector128 samples);
+
+ ///
+ /// Sums sixteen high-bit-depth reference samples.
+ ///
+ /// The reference samples.
+ /// The exact sum.
+ public static abstract int Sum(Vector256 samples);
+
+ ///
+ /// Sums thirty-two high-bit-depth reference samples.
+ ///
+ /// The reference samples.
+ /// The exact sum.
+ public static abstract int Sum(Vector512 samples);
+
+ ///
+ /// Calculates the 8-bit DC prediction.
+ ///
+ /// The sum of available reference samples.
+ /// The number of available reference samples.
+ /// The rounded DC prediction.
+ public static abstract byte Predict(int sum, int count);
+
+ ///
+ /// Calculates the high-bit-depth DC prediction.
+ ///
+ /// The sum of available reference samples.
+ /// The number of available reference samples.
+ /// The reconstructed sample precision.
+ /// The rounded DC prediction.
+ public static abstract short Predict(int sum, int count, int bitDepth);
+ }
+
+ ///
+ /// Predicts an 8-bit DC block.
+ ///
+ public static void Predict(bool hasLeft, bool hasAbove, Span destination, int destinationStride, ReadOnlySpan above, ReadOnlySpan left, int width, int height)
+ => Predictor.Predict(hasLeft, hasAbove, destination, destinationStride, above, left, width, height);
+
+ ///
+ /// Predicts a high-bit-depth DC block.
+ ///
+ public static void Predict(bool hasLeft, bool hasAbove, Span destination, int destinationStride, ReadOnlySpan above, ReadOnlySpan left, int width, int height, int bitDepth)
+ => Predictor.Predict(hasLeft, hasAbove, destination, destinationStride, above, left, width, height, bitDepth);
+
+ ///
+ /// Predicts an 8-bit DC block without hardware intrinsics.
+ ///
+ public static void PredictScalar(bool hasLeft, bool hasAbove, Span destination, int destinationStride, ReadOnlySpan above, ReadOnlySpan left, int width, int height)
+ => Predictor.PredictScalar(hasLeft, hasAbove, destination, destinationStride, above, left, width, height);
+
+ ///
+ /// Predicts a high-bit-depth DC block without hardware intrinsics.
+ ///
+ public static void PredictScalar(bool hasLeft, bool hasAbove, Span destination, int destinationStride, ReadOnlySpan above, ReadOnlySpan left, int width, int height, int bitDepth)
+ => Predictor.PredictScalar(hasLeft, hasAbove, destination, destinationStride, above, left, width, height, bitDepth);
+
+ ///
+ /// Calculates the DC value from the available neighboring samples.
+ ///
+ internal readonly struct DcOperator : IDcPredictionOperator
+ {
+ ///
+ [MethodImpl(MethodImplOptions.AggressiveInlining)]
+ public static int Sum(byte sample) => sample;
+
+ ///
+ [MethodImpl(MethodImplOptions.AggressiveInlining)]
+ public static int Sum(Vector128 samples)
+ {
+ (Vector128 lower, Vector128 upper) = Vector128.Widen(samples);
+
+ return Vector128.Sum(lower) + Vector128.Sum(upper);
+ }
+
+ ///
+ [MethodImpl(MethodImplOptions.AggressiveInlining)]
+ public static int Sum(Vector256 samples)
+ {
+ (Vector256 lower, Vector256 upper) = Vector256.Widen(samples);
+
+ return Vector256.Sum(lower) + Vector256.Sum(upper);
+ }
+
+ ///
+ [MethodImpl(MethodImplOptions.AggressiveInlining)]
+ public static int Sum(Vector512 samples)
+ {
+ (Vector512 lower, Vector512 upper) = Vector512.Widen(samples);
+
+ return Vector512.Sum(lower) + Vector512.Sum(upper);
+ }
+
+ ///
+ [MethodImpl(MethodImplOptions.AggressiveInlining)]
+ public static int Sum(short sample) => sample;
+
+ ///
+ [MethodImpl(MethodImplOptions.AggressiveInlining)]
+ public static int Sum(Vector128 samples)
+ {
+ (Vector128 lower, Vector128 upper) = Vector128.Widen(samples);
+
+ return Vector128.Sum(lower) + Vector128.Sum(upper);
+ }
+
+ ///
+ [MethodImpl(MethodImplOptions.AggressiveInlining)]
+ public static int Sum(Vector256 samples)
+ {
+ (Vector256 lower, Vector256 upper) = Vector256.Widen(samples);
+
+ return Vector256.Sum(lower) + Vector256.Sum(upper);
+ }
+
+ ///
+ [MethodImpl(MethodImplOptions.AggressiveInlining)]
+ public static int Sum(Vector512 samples)
+ {
+ (Vector512 lower, Vector512 upper) = Vector512.Widen(samples);
+
+ return Vector512.Sum(lower) + Vector512.Sum(upper);
+ }
+
+ ///
+ [MethodImpl(MethodImplOptions.AggressiveInlining)]
+ public static byte Predict(int sum, int count) => count == 0 ? (byte)128 : (byte)((sum + (count >> 1)) / count);
+
+ ///
+ [MethodImpl(MethodImplOptions.AggressiveInlining)]
+ public static short Predict(int sum, int count, int bitDepth)
+ => count == 0 ? (short)(1 << (bitDepth - 1)) : (short)((sum + (count >> 1)) / count);
+ }
+
+ ///
+ /// Reconstructs DC blocks through one closed reduction operator.
+ ///
+ /// The reference reduction and rounded mean arithmetic.
+ private static class Predictor
+ where TOperator : struct, IDcPredictionOperator
+ {
+ ///
+ /// Predicts an 8-bit DC block.
+ ///
+ /// Whether the left reference is available.
+ /// Whether the top reference is available.
+ /// The destination block.
+ /// The distance between destination rows.
+ /// The top reference.
+ /// The left reference.
+ /// The block width.
+ /// The block height.
+ public static void Predict(bool hasLeft, bool hasAbove, Span destination, int destinationStride, ReadOnlySpan above, ReadOnlySpan left, int width, int height)
+ {
+ int count = (hasAbove ? width : 0) + (hasLeft ? height : 0);
+ int sum = (hasAbove ? Sum(above[..width]) : 0) + (hasLeft ? Sum(left[..height]) : 0);
+ byte prediction = TOperator.Predict(sum, count);
+
+ for (int row = 0; row < height; row++)
+ {
+ destination.Slice(row * destinationStride, width).Fill(prediction);
+ }
+ }
+
+ ///
+ /// Predicts a high-bit-depth DC block.
+ ///
+ /// Whether the left reference is available.
+ /// Whether the top reference is available.
+ /// The destination block.
+ /// The distance between destination rows.
+ /// The top reference.
+ /// The left reference.
+ /// The block width.
+ /// The block height.
+ /// The reconstructed sample precision.
+ public static void Predict(bool hasLeft, bool hasAbove, Span destination, int destinationStride, ReadOnlySpan above, ReadOnlySpan left, int width, int height, int bitDepth)
+ {
+ int count = (hasAbove ? width : 0) + (hasLeft ? height : 0);
+ int sum = (hasAbove ? Sum(above[..width]) : 0) + (hasLeft ? Sum(left[..height]) : 0);
+ short prediction = TOperator.Predict(sum, count, bitDepth);
+
+ for (int row = 0; row < height; row++)
+ {
+ destination.Slice(row * destinationStride, width).Fill(prediction);
+ }
+ }
+
+ ///
+ /// Predicts an 8-bit DC block without hardware intrinsics.
+ ///
+ /// Whether the left reference is available.
+ /// Whether the top reference is available.
+ /// The destination block.
+ /// The distance between destination rows.
+ /// The top reference.
+ /// The left reference.
+ /// The block width.
+ /// The block height.
+ public static void PredictScalar(bool hasLeft, bool hasAbove, Span destination, int destinationStride, ReadOnlySpan above, ReadOnlySpan left, int width, int height)
+ {
+ int count = (hasAbove ? width : 0) + (hasLeft ? height : 0);
+ int sum = (hasAbove ? SumScalar(above[..width]) : 0) + (hasLeft ? SumScalar(left[..height]) : 0);
+ byte prediction = TOperator.Predict(sum, count);
+
+ for (int row = 0; row < height; row++)
+ {
+ ref byte destinationRow = ref destination[row * destinationStride];
+ for (int column = 0; column < width; column++)
+ {
+ Unsafe.Add(ref destinationRow, column) = prediction;
+ }
+ }
+ }
+
+ ///
+ /// Predicts a high-bit-depth DC block without hardware intrinsics.
+ ///
+ /// Whether the left reference is available.
+ /// Whether the top reference is available.
+ /// The destination block.
+ /// The distance between destination rows.
+ /// The top reference.
+ /// The left reference.
+ /// The block width.
+ /// The block height.
+ /// The reconstructed sample precision.
+ public static void PredictScalar(bool hasLeft, bool hasAbove, Span destination, int destinationStride, ReadOnlySpan above, ReadOnlySpan left, int width, int height, int bitDepth)
+ {
+ int count = (hasAbove ? width : 0) + (hasLeft ? height : 0);
+ int sum = (hasAbove ? SumScalar(above[..width]) : 0) + (hasLeft ? SumScalar(left[..height]) : 0);
+ short prediction = TOperator.Predict(sum, count, bitDepth);
+
+ for (int row = 0; row < height; row++)
+ {
+ ref short destinationRow = ref destination[row * destinationStride];
+ for (int column = 0; column < width; column++)
+ {
+ Unsafe.Add(ref destinationRow, column) = prediction;
+ }
+ }
+ }
+
+ ///
+ /// Sums 8-bit references through the widest available SIMD widths.
+ ///
+ /// The reference samples.
+ /// The exact sum.
+ private static int Sum(ReadOnlySpan samples)
+ {
+ ref byte samplesBase = ref MemoryMarshal.GetReference(samples);
+ int sum = 0;
+ int index = 0;
+
+ // The shared index deliberately continues through narrower widths. This handles every legal AV1 edge
+ // length without a separate dispatch tree and leaves only an incomplete final vector to scalar code.
+ if (Vector512.IsHardwareAccelerated)
+ {
+ int oneVectorFromEnd = samples.Length - Vector512.Count;
+ for (; index <= oneVectorFromEnd; index += Vector512.Count)
+ {
+ sum += TOperator.Sum(Vector512.LoadUnsafe(ref samplesBase, (nuint)index));
+ }
+ }
+
+ if (Vector256.IsHardwareAccelerated)
+ {
+ int oneVectorFromEnd = samples.Length - Vector256.Count;
+ for (; index <= oneVectorFromEnd; index += Vector256.Count)
+ {
+ sum += TOperator.Sum(Vector256.LoadUnsafe(ref samplesBase, (nuint)index));
+ }
+ }
+
+ if (Vector128.IsHardwareAccelerated)
+ {
+ int oneVectorFromEnd = samples.Length - Vector128.Count;
+ for (; index <= oneVectorFromEnd; index += Vector128.Count)
+ {
+ sum += TOperator.Sum(Vector128.LoadUnsafe(ref samplesBase, (nuint)index));
+ }
+ }
+
+ for (; index < samples.Length; index++)
+ {
+ sum += TOperator.Sum(Unsafe.Add(ref samplesBase, index));
+ }
+
+ return sum;
+ }
+
+ ///
+ /// Sums high-bit-depth references through the widest available SIMD widths.
+ ///
+ /// The reference samples.
+ /// The exact sum.
+ private static int Sum(ReadOnlySpan samples)
+ {
+ ref short samplesBase = ref MemoryMarshal.GetReference(samples);
+ int sum = 0;
+ int index = 0;
+
+ if (Vector512.IsHardwareAccelerated)
+ {
+ int oneVectorFromEnd = samples.Length - Vector512.Count;
+ for (; index <= oneVectorFromEnd; index += Vector512.Count)
+ {
+ sum += TOperator.Sum(Vector512.LoadUnsafe(ref samplesBase, (nuint)index));
+ }
+ }
+
+ if (Vector256.IsHardwareAccelerated)
+ {
+ int oneVectorFromEnd = samples.Length - Vector256.Count;
+ for (; index <= oneVectorFromEnd; index += Vector256.Count)
+ {
+ sum += TOperator.Sum(Vector256.LoadUnsafe(ref samplesBase, (nuint)index));
+ }
+ }
+
+ if (Vector128.IsHardwareAccelerated)
+ {
+ int oneVectorFromEnd = samples.Length - Vector128.Count;
+ for (; index <= oneVectorFromEnd; index += Vector128.Count)
+ {
+ sum += TOperator.Sum(Vector128.LoadUnsafe(ref samplesBase, (nuint)index));
+ }
+ }
+
+ for (; index < samples.Length; index++)
+ {
+ sum += TOperator.Sum(Unsafe.Add(ref samplesBase, index));
+ }
+
+ return sum;
+ }
+
+ ///
+ /// Sums 8-bit references without hardware intrinsics.
+ ///
+ /// The reference samples.
+ /// The exact sum.
+ private static int SumScalar(ReadOnlySpan samples)
+ {
+ ref byte samplesBase = ref MemoryMarshal.GetReference(samples);
+ int sum = 0;
+
+ for (int index = 0; index < samples.Length; index++)
+ {
+ sum += TOperator.Sum(Unsafe.Add(ref samplesBase, index));
+ }
+
+ return sum;
+ }
+
+ ///
+ /// Sums high-bit-depth references without hardware intrinsics.
+ ///
+ /// The reference samples.
+ /// The exact sum.
+ private static int SumScalar(ReadOnlySpan samples)
+ {
+ ref short samplesBase = ref MemoryMarshal.GetReference(samples);
+ int sum = 0;
+
+ for (int index = 0; index < samples.Length; index++)
+ {
+ sum += TOperator.Sum(Unsafe.Add(ref samplesBase, index));
+ }
+
+ return sum;
+ }
+ }
+}
diff --git a/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1DcIntraPredictor.cs b/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1DcIntraPredictor.cs
deleted file mode 100644
index 7168c56b4..000000000
--- a/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1DcIntraPredictor.cs
+++ /dev/null
@@ -1,254 +0,0 @@
-// Copyright (c) Six Labors.
-// Licensed under the Six Labors Split License.
-
-using System.Runtime.CompilerServices;
-using System.Runtime.InteropServices;
-using System.Runtime.Intrinsics;
-
-namespace SixLabors.ImageSharp.Formats.Heif.Av1.Prediction;
-
-///
-/// Reconstructs AV1 DC intra-prediction blocks from the available neighboring samples.
-///
-///
-/// Reference reduction uses the widest available integer lanes, then one rounded scalar mean is broadcast across each
-/// destination row. Row filling is delegated to span operations so the runtime supplies its native vectorized store;
-/// explicit scalar entry points remain available for feature-disabled conformance tests.
-///
-internal static class Av1DcIntraPredictor
-{
- ///
- /// Predicts an 8-bit DC block.
- ///
- /// Whether the prepared left reference is available.
- /// Whether the prepared top reference is available.
- /// The destination block origin.
- /// The destination row stride in samples.
- /// The prepared top reference.
- /// The prepared left reference.
- /// The block width in samples.
- /// The block height in samples.
- public static void Predict(bool hasLeft, bool hasAbove, Span destination, int destinationStride, ReadOnlySpan above, ReadOnlySpan left, int width, int height)
- {
- int count = (hasAbove ? width : 0) + (hasLeft ? height : 0);
- int sum = (hasAbove ? Sum(above[..width]) : 0) + (hasLeft ? Sum(left[..height]) : 0);
-
- // Section 7.11.2.2 defines the unsigned midpoint when no reference is available. Otherwise adding half
- // the reference count implements the specified rounded mean before extending it across the whole block.
- byte prediction = count == 0 ? (byte)128 : (byte)((sum + (count >> 1)) / count);
- for (int row = 0; row < height; row++)
- {
- destination.Slice(row * destinationStride, width).Fill(prediction);
- }
- }
-
- ///
- /// Predicts a high-bit-depth DC block.
- ///
- /// Whether the prepared left reference is available.
- /// Whether the prepared top reference is available.
- /// The destination block origin.
- /// The destination row stride in samples.
- /// The prepared top reference.
- /// The prepared left reference.
- /// The block width in samples.
- /// The block height in samples.
- /// The reconstructed sample precision.
- public static void Predict(bool hasLeft, bool hasAbove, Span destination, int destinationStride, ReadOnlySpan above, ReadOnlySpan left, int width, int height, int bitDepth)
- {
- int count = (hasAbove ? width : 0) + (hasLeft ? height : 0);
- int sum = (hasAbove ? Sum(above[..width]) : 0) + (hasLeft ? Sum(left[..height]) : 0);
- short prediction = count == 0 ? (short)(1 << (bitDepth - 1)) : (short)((sum + (count >> 1)) / count);
- for (int row = 0; row < height; row++)
- {
- destination.Slice(row * destinationStride, width).Fill(prediction);
- }
- }
-
- ///
- /// Predicts an 8-bit DC block without hardware intrinsics.
- ///
- /// Whether the prepared left reference is available.
- /// Whether the prepared top reference is available.
- /// The destination block origin.
- /// The destination row stride in samples.
- /// The prepared top reference.
- /// The prepared left reference.
- /// The block width in samples.
- /// The block height in samples.
- public static void PredictScalar(bool hasLeft, bool hasAbove, Span destination, int destinationStride, ReadOnlySpan above, ReadOnlySpan left, int width, int height)
- {
- int count = (hasAbove ? width : 0) + (hasLeft ? height : 0);
- int sum = (hasAbove ? SumScalar(above[..width]) : 0) + (hasLeft ? SumScalar(left[..height]) : 0);
- byte prediction = count == 0 ? (byte)128 : (byte)((sum + (count >> 1)) / count);
- for (int row = 0; row < height; row++)
- {
- ref byte destinationRow = ref destination[row * destinationStride];
- for (int column = 0; column < width; column++)
- {
- Unsafe.Add(ref destinationRow, column) = prediction;
- }
- }
- }
-
- ///
- /// Predicts a high-bit-depth DC block without hardware intrinsics.
- ///
- /// Whether the prepared left reference is available.
- /// Whether the prepared top reference is available.
- /// The destination block origin.
- /// The destination row stride in samples.
- /// The prepared top reference.
- /// The prepared left reference.
- /// The block width in samples.
- /// The block height in samples.
- /// The reconstructed sample precision.
- public static void PredictScalar(bool hasLeft, bool hasAbove, Span destination, int destinationStride, ReadOnlySpan above, ReadOnlySpan left, int width, int height, int bitDepth)
- {
- int count = (hasAbove ? width : 0) + (hasLeft ? height : 0);
- int sum = (hasAbove ? SumScalar(above[..width]) : 0) + (hasLeft ? SumScalar(left[..height]) : 0);
- short prediction = count == 0 ? (short)(1 << (bitDepth - 1)) : (short)((sum + (count >> 1)) / count);
- for (int row = 0; row < height; row++)
- {
- ref short destinationRow = ref destination[row * destinationStride];
- for (int column = 0; column < width; column++)
- {
- Unsafe.Add(ref destinationRow, column) = prediction;
- }
- }
- }
-
- ///
- /// Sums 8-bit neighboring samples using the widest available SIMD width.
- ///
- /// The samples to sum.
- /// The exact sum.
- private static int Sum(ReadOnlySpan samples)
- {
- ref byte samplesBase = ref MemoryMarshal.GetReference(samples);
- int sum = 0;
- int index = 0;
-
- // Widening prevents the packed-byte reduction from overflowing before the horizontal sum. The shared index
- // advances through every available vector width and leaves only the incomplete final group to the scalar loop.
- if (Vector512.IsHardwareAccelerated)
- {
- int oneVectorFromEnd = samples.Length - Vector512.Count;
- for (; index <= oneVectorFromEnd; index += Vector512.Count)
- {
- (Vector512 low, Vector512 high) = Vector512.Widen(Vector512.LoadUnsafe(ref samplesBase, (nuint)index));
- sum += Vector512.Sum(low) + Vector512.Sum(high);
- }
- }
-
- if (Vector256.IsHardwareAccelerated)
- {
- int oneVectorFromEnd = samples.Length - Vector256.Count;
- for (; index <= oneVectorFromEnd; index += Vector256.Count)
- {
- (Vector256 low, Vector256 high) = Vector256.Widen(Vector256.LoadUnsafe(ref samplesBase, (nuint)index));
- sum += Vector256.Sum(low) + Vector256.Sum(high);
- }
- }
-
- if (Vector128.IsHardwareAccelerated)
- {
- int oneVectorFromEnd = samples.Length - Vector128.Count;
- for (; index <= oneVectorFromEnd; index += Vector128.Count)
- {
- (Vector128 low, Vector128 high) = Vector128.Widen(Vector128.LoadUnsafe(ref samplesBase, (nuint)index));
- sum += Vector128.Sum(low) + Vector128.Sum(high);
- }
- }
-
- for (; index < samples.Length; index++)
- {
- sum += Unsafe.Add(ref samplesBase, index);
- }
-
- return sum;
- }
-
- ///
- /// Sums high-bit-depth neighboring samples using the widest available SIMD width.
- ///
- /// The samples to sum.
- /// The exact sum.
- private static int Sum(ReadOnlySpan samples)
- {
- ref short samplesBase = ref MemoryMarshal.GetReference(samples);
- int sum = 0;
- int index = 0;
-
- // Signed 16-bit storage is nonnegative for supported bit depths. Widening to Int32 preserves the exact sum of
- // the largest permitted reference edge before the rounded mean is calculated once outside this loop.
- if (Vector512.IsHardwareAccelerated)
- {
- int oneVectorFromEnd = samples.Length - Vector512.Count;
- for (; index <= oneVectorFromEnd; index += Vector512.Count)
- {
- (Vector512 low, Vector512 high) = Vector512.Widen(Vector512.LoadUnsafe(ref samplesBase, (nuint)index));
- sum += Vector512.Sum(low) + Vector512.Sum(high);
- }
- }
-
- if (Vector256.IsHardwareAccelerated)
- {
- int oneVectorFromEnd = samples.Length - Vector256.Count;
- for (; index <= oneVectorFromEnd; index += Vector256.Count)
- {
- (Vector256 low, Vector256 high) = Vector256.Widen(Vector256.LoadUnsafe(ref samplesBase, (nuint)index));
- sum += Vector256.Sum(low) + Vector256.Sum(high);
- }
- }
-
- if (Vector128.IsHardwareAccelerated)
- {
- int oneVectorFromEnd = samples.Length - Vector128.Count;
- for (; index <= oneVectorFromEnd; index += Vector128.Count)
- {
- (Vector128 low, Vector128 high) = Vector128.Widen(Vector128.LoadUnsafe(ref samplesBase, (nuint)index));
- sum += Vector128.Sum(low) + Vector128.Sum(high);
- }
- }
-
- for (; index < samples.Length; index++)
- {
- sum += Unsafe.Add(ref samplesBase, index);
- }
-
- return sum;
- }
-
- ///
- /// Sums 8-bit neighboring samples without hardware intrinsics.
- ///
- /// The samples to sum.
- /// The exact sum.
- private static int SumScalar(ReadOnlySpan samples)
- {
- int sum = 0;
- foreach (byte sample in samples)
- {
- sum += sample;
- }
-
- return sum;
- }
-
- ///
- /// Sums high-bit-depth neighboring samples without hardware intrinsics.
- ///
- /// The samples to sum.
- /// The exact sum.
- private static int SumScalar(ReadOnlySpan samples)
- {
- int sum = 0;
- foreach (short sample in samples)
- {
- sum += sample;
- }
-
- return sum;
- }
-}
diff --git a/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1DirectionalIntraPredictor.Operations.cs b/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1DirectionalIntraPredictor.Operations.cs
index 26b8d8834..4c5f7e7fc 100644
--- a/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1DirectionalIntraPredictor.Operations.cs
+++ b/src/ImageSharp/Formats/Heif/Av1/Prediction/Av1DirectionalIntraPredictor.Operations.cs
@@ -4,6 +4,7 @@
using System.Runtime.CompilerServices;
using System.Runtime.InteropServices;
using System.Runtime.Intrinsics;
+using SixLabors.ImageSharp.Formats.Heif.Av1.Transform;
namespace SixLabors.ImageSharp.Formats.Heif.Av1.Prediction;
@@ -17,615 +18,933 @@ namespace SixLabors.ImageSharp.Formats.Heif.Av1.Prediction;
internal static partial class Av1DirectionalIntraPredictor
{
///
- /// Gets the indices of even bytes in one upsampled reference vector.
+ /// Implements the directional traversal for one closed interpolation operator.
///
- private static Vector128 EvenByteIndices => Vector128.Create((byte)0, 2, 4, 6, 8, 10, 12, 14, 0, 2, 4, 6, 8, 10, 12, 14);
-
- ///
- /// Gets the indices of odd bytes in one upsampled reference vector.
- ///
- private static Vector128 OddByteIndices => Vector128.Create((byte)1, 3, 5, 7, 9, 11, 13, 15, 1, 3, 5, 7, 9, 11, 13, 15);
-
- ///
- /// Gets the indices of even high-bit-depth samples in one upsampled reference vector.
- ///
- private static Vector128 EvenShortIndices => Vector128.Create((short)0, 2, 4, 6, 0, 2, 4, 6);
-
- ///
- /// Gets the indices of odd high-bit-depth samples in one upsampled reference vector.
- ///
- private static Vector128 OddShortIndices => Vector128.Create((short)1, 3, 5, 7, 1, 3, 5, 7);
-
- ///
- /// Predicts one 8-bit zone 1 block using contiguous SIMD projection rows.
- ///
- /// The destination block origin.
- /// The destination row stride.
- /// The projected top reference.
- /// Whether the reference contains half-sample positions.
- /// The Q8 projection derivative.
- /// The block width.
- /// The block height.
- private static void PredictZone1(Span destination, int destinationStride, ReadOnlySpan above, bool upsample, int derivative, int width, int height)
+ private static partial class Predictor
+ where TOperator : struct, IDirectionalPredictionOperator
{
- int upsampleShift = upsample ? 1 : 0;
- int maximumBasis = (width + height - 1) << upsampleShift;
- int fractionBits = 6 - upsampleShift;
- int projection = derivative;
-
- for (int row = 0; row < height; row++, projection += derivative)
+ ///
+ /// Gets the indices of even bytes in one upsampled reference vector.
+ ///
+ private static Vector128 EvenByteIndices => Vector128.Create((byte)0, 2, 4, 6, 8, 10, 12, 14, 0, 2, 4, 6, 8, 10, 12, 14);
+
+ ///
+ /// Gets the indices of odd bytes in one upsampled reference vector.
+ ///
+ private static Vector128 OddByteIndices => Vector128.Create((byte)1, 3, 5, 7, 9, 11, 13, 15, 1, 3, 5, 7, 9, 11, 13, 15);
+
+ ///
+ /// Gets the indices of even high-bit-depth samples in one upsampled reference vector.
+ ///
+ private static Vector128 EvenShortIndices => Vector128.Create((short)0, 2, 4, 6, 0, 2, 4, 6);
+
+ ///
+ /// Gets the indices of odd high-bit-depth samples in one upsampled reference vector.
+ ///
+ private static Vector128 OddShortIndices => Vector128.Create((short)1, 3, 5, 7, 1, 3, 5, 7);
+
+ ///
+ /// Predicts one 8-bit zone 1 block using contiguous SIMD projection rows.
+ ///
+ /// The destination block origin.
+ /// The destination row stride.
+ /// The projected top reference.
+ /// Whether the reference contains half-sample positions.
+ /// The Q8 projection derivative.
+ /// The block width.
+ /// The block height.
+ private static void PredictZone1(Span destination, int destinationStride, ReadOnlySpan above, bool upsample, int derivative, int width, int height)
{
- int basis = projection >> fractionBits;
- int weight = ((projection << upsampleShift) & 0x3F) >> 1;
- InterpolateRow(destination.Slice(row * destinationStride, width), above, basis, weight, upsample, maximumBasis);
- }
- }
+ int upsampleShift = upsample ? 1 : 0;
+ int maximumBasis = (width + height - 1) << upsampleShift;
+ int fractionBits = 6 - upsampleShift;
+ int projection = derivative;
- ///
- /// Predicts one high-bit-depth zone 1 block using contiguous SIMD projection rows.
- ///
- /// The destination block origin.
- /// The destination row stride.
- /// The projected top reference.
- /// Whether the reference contains half-sample positions.
- /// The Q8 projection derivative.
- /// The block width.
- /// The block height.
- private static void PredictZone1(Span destination, int destinationStride, ReadOnlySpan above, bool upsample, int derivative, int width, int height)
- {
- int upsampleShift = upsample ? 1 : 0;
- int maximumBasis = (width + height - 1) << upsampleShift;
- int fractionBits = 6 - upsampleShift;
- int projection = derivative;
-
- for (int row = 0; row < height; row++, projection += derivative)
- {
- int basis = projection >> fractionBits;
- int weight = ((projection << upsampleShift) & 0x3F) >> 1;
- InterpolateRow(destination.Slice(row * destinationStride, width), above, basis, weight, upsample, maximumBasis);
+ for (int row = 0; row < height; row++, projection += derivative)
+ {
+ int basis = projection >> fractionBits;
+ int weight = ((projection << upsampleShift) & 0x3F) >> 1;
+ InterpolateRow(destination.Slice(row * destinationStride, width), above, basis, weight, upsample, maximumBasis);
+ }
}
- }
- ///
- /// Predicts one 8-bit zone 2 block using vectorized left gathers and contiguous top projection rows.
- ///
- /// The destination block origin.
- /// The destination row stride.
- /// The projected top reference.
- /// The projected left reference.
- /// Whether the top reference contains half-sample positions.
- /// Whether the left reference contains half-sample positions.
- /// The horizontal Q8 derivative.
- /// The vertical Q8 derivative.
- /// The block width.
- /// The block height.
- private static void PredictZone2(Span destination, int destinationStride, ReadOnlySpan above, ReadOnlySpan left, bool upsampleAbove, bool upsampleLeft, int dx, int dy, int width, int height)
- {
- int aboveShift = upsampleAbove ? 1 : 0;
- int minimumTopBasis = -(1 << aboveShift);
- int topFractionBits = 6 - aboveShift;
- int topBasisIncrement = 1 << aboveShift;
- int topProjection = -dx;
-
- for (int row = 0; row < height; row++, topProjection -= dx)
+ ///
+ /// Predicts one high-bit-depth zone 1 block using contiguous SIMD projection rows.
+ ///
+ /// The destination block origin.
+ /// The destination row stride.
+ /// The projected top reference.
+ /// Whether the reference contains half-sample positions.
+ /// The Q8 projection derivative.
+ /// The block width.
+ /// The block height.
+ private static void PredictZone1(Span destination, int destinationStride, ReadOnlySpan above, bool upsample, int derivative, int width, int height)
{
- int topBasis = topProjection >> topFractionBits;
- int leftCount = 0;
- while (leftCount < width && topBasis < minimumTopBasis)
+ int upsampleShift = upsample ? 1 : 0;
+ int maximumBasis = (width + height - 1) << upsampleShift;
+ int fractionBits = 6 - upsampleShift;
+ int projection = derivative;
+
+ for (int row = 0; row < height; row++, projection += derivative)
{
- leftCount++;
- topBasis += topBasisIncrement;
+ int basis = projection >> fractionBits;
+ int weight = ((projection << upsampleShift) & 0x3F) >> 1;
+ InterpolateRow(destination.Slice(row * destinationStride, width), above, basis, weight, upsample, maximumBasis);
}
+ }
- Span destinationRow = destination.Slice(row * destinationStride, width);
- int leftProjection = (row << 6) - dy;
- InterpolateLeft(destinationRow[..leftCount], left, leftProjection, dy, upsampleLeft);
+ ///
+ /// Predicts one 8-bit zone 2 block using vectorized left gathers and contiguous top projection rows.
+ ///
+ /// The destination block origin.
+ /// The destination row stride.
+ /// The projected top reference.
+ /// The projected left reference.
+ /// Whether the top reference contains half-sample positions.
+ /// Whether the left reference contains half-sample positions.
+ /// The horizontal Q8 derivative.
+ /// The vertical Q8 derivative.
+ /// The block width.
+ /// The block height.
+ private static void PredictZone2(Span destination, int destinationStride, ReadOnlySpan above, ReadOnlySpan left, bool upsampleAbove, bool upsampleLeft, int dx, int dy, int width, int height)
+ {
+ int aboveShift = upsampleAbove ? 1 : 0;
+ int minimumTopBasis = -(1 << aboveShift);
+ int topFractionBits = 6 - aboveShift;
+ int topBasisIncrement = 1 << aboveShift;
+ int topProjection = -dx;
- if (leftCount < width)
+ for (int row = 0; row < height; row++, topProjection -= dx)
{
- int topWeight = ((topProjection << aboveShift) & 0x3F) >> 1;
- InterpolateRow(destinationRow[leftCount..], above, topBasis, topWeight, upsampleAbove, int.MaxValue);
+ int topBasis = topProjection >> topFractionBits;
+ int leftCount = 0;
+ while (leftCount < width && topBasis < minimumTopBasis)
+ {
+ leftCount++;
+ topBasis += topBasisIncrement;
+ }
+
+ Span destinationRow = destination.Slice(row * destinationStride, width);
+ int leftProjection = (row << 6) - dy;
+ InterpolateLeft(destinationRow[..leftCount], left, leftProjection, dy, upsampleLeft);
+
+ if (leftCount < width)
+ {
+ int topWeight = ((topProjection << aboveShift) & 0x3F) >> 1;
+ InterpolateRow(destinationRow[leftCount..], above, topBasis, topWeight, upsampleAbove, int.MaxValue);
+ }
}
}
- }
-
- ///
- /// Predicts one high-bit-depth zone 2 block using vectorized left gathers and contiguous top projection rows.
- ///
- /// The destination block origin.
- /// The destination row stride.
- /// The projected top reference.
- /// The projected left reference.
- /// Whether the top reference contains half-sample positions.
- /// Whether the left reference contains half-sample positions.
- /// The horizontal Q8 derivative.
- /// The vertical Q8 derivative.
- /// The block width.
- /// The block height.
- private static void PredictZone2(Span destination, int destinationStride, ReadOnlySpan above, ReadOnlySpan left, bool upsampleAbove, bool upsampleLeft, int dx, int dy, int width, int height)
- {
- int aboveShift = upsampleAbove ? 1 : 0;
- int minimumTopBasis = -(1 << aboveShift);
- int topFractionBits = 6 - aboveShift;
- int topBasisIncrement = 1 << aboveShift;
- int topProjection = -dx;
- for (int row = 0; row < height; row++, topProjection -= dx)
+ ///
+ /// Predicts one high-bit-depth zone 2 block using vectorized left gathers and contiguous top projection rows.
+ ///
+ /// The destination block origin.
+ /// The destination row stride.
+ /// The projected top reference.
+ /// The projected left reference.
+ /// Whether the top reference contains half-sample positions.
+ /// Whether the left reference contains half-sample positions.
+ /// The horizontal Q8 derivative.
+ /// The vertical Q8 derivative.
+ /// The block width.
+ /// The block height.
+ private static void PredictZone2(Span destination, int destinationStride, ReadOnlySpan above, ReadOnlySpan left, bool upsampleAbove, bool upsampleLeft, int dx, int dy, int width, int height)
{
- int topBasis = topProjection >> topFractionBits;
- int leftCount = 0;
- while (leftCount < width && topBasis < minimumTopBasis)
+ int aboveShift = upsampleAbove ? 1 : 0;
+ int minimumTopBasis = -(1 << aboveShift);
+ int topFractionBits = 6 - aboveShift;
+ int topBasisIncrement = 1 << aboveShift;
+ int topProjection = -dx;
+
+ for (int row = 0; row < height; row++, topProjection -= dx)
{
- leftCount++;
- topBasis += topBasisIncrement;
- }
+ int topBasis = topProjection >> topFractionBits;
+ int leftCount = 0;
+ while (leftCount < width && topBasis < minimumTopBasis)
+ {
+ leftCount++;
+ topBasis += topBasisIncrement;
+ }
- Span destinationRow = destination.Slice(row * destinationStride, width);
- int leftProjection = (row << 6) - dy;
- InterpolateLeft(destinationRow[..leftCount], left, leftProjection, dy, upsampleLeft);
+ Span destinationRow = destination.Slice(row * destinationStride, width);
+ int leftProjection = (row << 6) - dy;
+ InterpolateLeft(destinationRow[..leftCount], left, leftProjection, dy, upsampleLeft);
- if (leftCount < width)
- {
- int topWeight = ((topProjection << aboveShift) & 0x3F) >> 1;
- InterpolateRow(destinationRow[leftCount..], above, topBasis, topWeight, upsampleAbove, int.MaxValue);
+ if (leftCount < width)
+ {
+ int topWeight = ((topProjection << aboveShift) & 0x3F) >> 1;
+ InterpolateRow(destinationRow[leftCount..], above, topBasis, topWeight, upsampleAbove, int.MaxValue);
+ }
}
}
- }
- ///
- /// Interpolates one 8-bit projection row.
- ///
- /// The destination row.
- /// The projected reference samples.
- /// The first integral reference coordinate.
- /// The right-sample interpolation weight.
- /// Whether consecutive output samples advance two reference positions.
- /// The final extended reference coordinate, or when the row cannot reach it.
- private static void InterpolateRow(Span destination, ReadOnlySpan reference, int basis, int weight, bool upsample, int maximumBasis)
- {
- ref byte destinationBase = ref MemoryMarshal.GetReference(destination);
- ref byte referenceBase = ref MemoryMarshal.GetReference(reference);
- int basisIncrement = upsample ? 2 : 1;
- int validCount = maximumBasis == int.MaxValue || basis >= maximumBasis
- ? maximumBasis == int.MaxValue ? destination.Length : 0
- : Math.Min(destination.Length, ((maximumBasis - 1 - basis) / basisIncrement) + 1);
- int index = 0;
-
- if (!upsample)
+ ///
+ /// Interpolates one 8-bit projection row.
+ ///
+ /// The destination row.
+ /// The projected reference samples.
+ /// The first integral reference coordinate.
+ /// The right-sample interpolation weight.
+ /// Whether consecutive output samples advance two reference positions.
+ /// The final extended reference coordinate, or when the row cannot reach it.
+ private static void InterpolateRow(Span destination, ReadOnlySpan reference, int basis, int weight, bool upsample, int maximumBasis)
{
- // A single index is advanced through all supported widths. A narrower path consumes only the remainder
- // left by the wider path, so the row is written once without requiring padded destination storage.
- if (Vector512.IsHardwareAccelerated)
+ ref byte destinationBase = ref MemoryMarshal.GetReference(destination);
+ ref byte referenceBase = ref MemoryMarshal.GetReference(reference);
+ int basisIncrement = upsample ? 2 : 1;
+ int validCount = maximumBasis == int.MaxValue || basis >= maximumBasis
+ ? maximumBasis == int.MaxValue ? destination.Length : 0
+ : Math.Min(destination.Length, ((maximumBasis - 1 - basis) / basisIncrement) + 1);
+ int index = 0;
+
+ if (!upsample)
{
- int oneVectorFromEnd = validCount - Vector512.Count;
- for (; index <= oneVectorFromEnd; index += Vector512.Count)
+ // A single index is advanced through all supported widths. A narrower path consumes only the remainder
+ // left by the wider path, so the row is written once without requiring padded destination storage.
+ if (Vector512.IsHardwareAccelerated)
{
- Vector512 left = Vector512.LoadUnsafe(ref referenceBase, (nuint)(basis + index));
- Vector512 right = Vector512.LoadUnsafe(ref referenceBase, (nuint)(basis + index + 1));
- Interpolate(left, right, weight).StoreUnsafe(ref destinationBase, (nuint)index);
+ int oneVectorFromEnd = validCount - Vector512.Count;
+ for (; index <= oneVectorFromEnd; index += Vector512.Count)
+ {
+ Vector512 left = Vector512.LoadUnsafe(ref referenceBase, (nuint)(basis + index));
+ Vector512 right = Vector512.LoadUnsafe(ref referenceBase, (nuint)(basis + index + 1));
+ TOperator.Interpolate(left, right, weight).StoreUnsafe(ref destinationBase, (nuint)index);
+ }
+ }
+
+ if (Vector256.IsHardwareAccelerated)
+ {
+ int oneVectorFromEnd = validCount - Vector256.Count;
+ for (; index <= oneVectorFromEnd; index += Vector256.Count)
+ {
+ Vector256 left = Vector256.LoadUnsafe(ref referenceBase, (nuint)(basis + index));
+ Vector256 right = Vector256.LoadUnsafe(ref referenceBase, (nuint)(basis + index + 1));
+ TOperator.Interpolate(left, right, weight).StoreUnsafe(ref destinationBase, (nuint)index);
+ }
}
- }
- if (Vector256.IsHardwareAccelerated)
+ if (Vector128.IsHardwareAccelerated)
+ {
+ int oneVectorFromEnd = validCount - Vector128.Count;
+ for (; index <= oneVectorFromEnd; index += Vector128.Count)
+ {
+ Vector128 left = Vector128.LoadUnsafe(ref referenceBase, (nuint)(basis + index));
+ Vector128 right = Vector128.LoadUnsafe(ref referenceBase, (nuint)(basis + index + 1));
+ TOperator.Interpolate(left, right, weight).StoreUnsafe(ref destinationBase, (nuint)index);
+ }
+ }
+ }
+ else if (Vector128.IsHardwareAccelerated)
{
- int oneVectorFromEnd = validCount - Vector256.Count;
- for (; index <= oneVectorFromEnd; index += Vector256.Count)
+ int oneVectorFromEnd = validCount - 8;
+ for (; index <= oneVectorFromEnd; index += 8)
{
- Vector256 left = Vector256.LoadUnsafe(ref referenceBase, (nuint)(basis + index));
- Vector256 right = Vector256.LoadUnsafe(ref referenceBase, (nuint)(basis + index + 1));
- Interpolate(left, right, weight).StoreUnsafe(ref destinationBase, (nuint)index);
+ Vector128 source = Vector128.LoadUnsafe(ref referenceBase, (nuint)(basis + (index * 2)));
+
+ // ShuffleNative maps to byte-table lookup on AdvSimd and PSHUFB on x86. Eight output samples are
+ // gathered from sixteen half-sample positions without scalar lane construction.
+ Vector128 left = Vector128.ShuffleNative(source, EvenByteIndices);
+ Vector128 right = Vector128.ShuffleNative(source, OddByteIndices);
+ Vector128 prediction = TOperator.Interpolate(left, right, weight);
+ Unsafe.As(ref Unsafe.Add(ref destinationBase, index)) = prediction.AsUInt64().ToScalar();
}
}
if (Vector128.IsHardwareAccelerated)
{
- int oneVectorFromEnd = validCount - Vector128.Count;
- for (; index <= oneVectorFromEnd; index += Vector128.Count)
+ // Four-lane construction covers both the final non-upsampled remainder and targets without a native
+ // gather. Each lane carries an independently projected coordinate but shares the row's interpolation weight.
+ int oneVectorFromEnd = validCount - 4;
+ for (; index <= oneVectorFromEnd; index += 4)
{
- Vector128 left = Vector128.LoadUnsafe(ref referenceBase, (nuint)(basis + index));
- Vector128 right = Vector128.LoadUnsafe(ref referenceBase, (nuint)(basis + index + 1));
- Interpolate(left, right, weight).StoreUnsafe(ref destinationBase, (nuint)index);
+ int source = basis + (index * basisIncrement);
+ Vector128 left = Vector128.Create(
+ (int)Unsafe.Add(ref referenceBase, source),
+ Unsafe.Add(ref referenceBase, source + basisIncrement),
+ Unsafe.Add(ref referenceBase, source + (2 * basisIncrement)),
+ Unsafe.Add(ref referenceBase, source + (3 * basisIncrement)));
+
+ Vector128 right = Vector128.Create(
+ (int)Unsafe.Add(ref referenceBase, source + 1),
+ Unsafe.Add(ref referenceBase, source + basisIncrement + 1),
+ Unsafe.Add(ref referenceBase, source + (2 * basisIncrement) + 1),
+ Unsafe.Add(ref referenceBase, source + (3 * basisIncrement) + 1));
+
+ StoreFourBytes(TOperator.Interpolate(left, right, Vector128.Create(weight)), ref Unsafe.Add(ref destinationBase, index));
}
}
- }
- else if (Vector128.IsHardwareAccelerated)
- {
- int oneVectorFromEnd = validCount - 8;
- for (; index <= oneVectorFromEnd; index += 8)
- {
- Vector128 source = Vector128.LoadUnsafe(ref referenceBase, (nuint)(basis + (index * 2)));
-
- // ShuffleNative maps to byte-table lookup on AdvSimd and PSHUFB on x86. Eight output samples are
- // gathered from sixteen half-sample positions without scalar lane construction.
- Vector128 left = Vector128.ShuffleNative(source, EvenByteIndices);
- Vector128 right = Vector128.ShuffleNative(source, OddByteIndices);
- Vector128 prediction = Interpolate(left, right, weight);
- Unsafe.As(ref Unsafe.Add(ref destinationBase, index)) = prediction.AsUInt64().GetElement(0);
- }
- }
- if (Vector128.IsHardwareAccelerated)
- {
- // Four-lane construction covers both the final non-upsampled remainder and targets without a native
- // gather. Each lane carries an independently projected coordinate but shares the row's interpolation weight.
- int oneVectorFromEnd = validCount - 4;
- for (; index <= oneVectorFromEnd; index += 4)
+ for (; index < validCount; index++)
{
int source = basis + (index * basisIncrement);
- Vector128 left = Vector128.Create(
- (int)Unsafe.Add(ref referenceBase, source),
- Unsafe.Add(ref referenceBase, source + basisIncrement),
- Unsafe.Add(ref referenceBase, source + (2 * basisIncrement)),
- Unsafe.Add(ref referenceBase, source + (3 * basisIncrement)));
-
- Vector128 right = Vector128.Create(
- (int)Unsafe.Add(ref referenceBase, source + 1),
- Unsafe.Add(ref referenceBase, source + basisIncrement + 1),
- Unsafe.Add(ref referenceBase, source + (2 * basisIncrement) + 1),
- Unsafe.Add(ref referenceBase, source + (3 * basisIncrement) + 1));
+ Unsafe.Add(ref destinationBase, index) = (byte)(((Unsafe.Add(ref referenceBase, source) * (32 - weight)) + (Unsafe.Add(ref referenceBase, source + 1) * weight) + 16) >> 5);
+ }
- StoreFourBytes(Interpolate(left, right, Vector128.Create(weight)), ref Unsafe.Add(ref destinationBase, index));
+ if (validCount < destination.Length)
+ {
+ destination[validCount..].Fill(Unsafe.Add(ref referenceBase, maximumBasis));
}
}
- for (; index < validCount; index++)
+ ///
+ /// Interpolates one high-bit-depth projection row.
+ ///
+ /// The destination row.
+ /// The projected reference samples.
+ /// The first integral reference coordinate.
+ /// The right-sample interpolation weight.
+ /// Whether consecutive output samples advance two reference positions.
+ /// The final extended reference coordinate, or when the row cannot reach it.
+ private static void InterpolateRow(Span destination, ReadOnlySpan reference, int basis, int weight, bool upsample, int maximumBasis)
{
- int source = basis + (index * basisIncrement);
- Unsafe.Add(ref destinationBase, index) = (byte)(((Unsafe.Add(ref referenceBase, source) * (32 - weight)) + (Unsafe.Add(ref referenceBase, source + 1) * weight) + 16) >> 5);
- }
+ ref short destinationBase = ref MemoryMarshal.GetReference(destination);
+ ref short referenceBase = ref MemoryMarshal.GetReference(reference);
+ int basisIncrement = upsample ? 2 : 1;
+ int validCount = maximumBasis == int.MaxValue || basis >= maximumBasis
+ ? maximumBasis == int.MaxValue ? destination.Length : 0
+ : Math.Min(destination.Length, ((maximumBasis - 1 - basis) / basisIncrement) + 1);
+ int index = 0;
+
+ if (!upsample)
+ {
+ // High-bit-depth samples stay in signed 16-bit storage, but interpolation widens to Int32 before the Q5
+ // weighted sum. The largest supported 12-bit sample therefore cannot overflow an intermediate lane.
+ if (Vector512.IsHardwareAccelerated)
+ {
+ int oneVectorFromEnd = validCount - Vector512.Count;
+ for (; index <= oneVectorFromEnd; index += Vector512.Count)
+ {
+ Vector512 left = Vector512.LoadUnsafe(ref referenceBase, (nuint)(basis + index));
+ Vector512 right = Vector512.LoadUnsafe(ref referenceBase, (nuint)(basis + index + 1));
+ TOperator.Interpolate(left, right, weight).StoreUnsafe(ref destinationBase, (nuint)index);
+ }
+ }
- if (validCount < destination.Length)
- {
- destination[validCount..].Fill(Unsafe.Add(ref referenceBase, maximumBasis));
- }
- }
+ if (Vector256.IsHardwareAccelerated)
+ {
+ int oneVectorFromEnd = validCount - Vector256.Count;
+ for (; index <= oneVectorFromEnd; index += Vector256.Count)
+ {
+ Vector256 left = Vector256.LoadUnsafe(ref referenceBase, (nuint)(basis + index));
+ Vector256 right = Vector256.LoadUnsafe(ref referenceBase, (nuint)(basis + index + 1));
+ TOperator.Interpolate(left, right, weight).StoreUnsafe(ref destinationBase, (nuint)index);
+ }
+ }
- ///
- /// Interpolates one high-bit-depth projection row.
- ///
- /// The destination row.
- /// The projected reference samples.
- /// The first integral reference coordinate.
- /// The right-sample interpolation weight.
- /// Whether consecutive output samples advance two reference positions.
- /// The final extended reference coordinate, or when the row cannot reach it.
- private static void InterpolateRow(Span destination, ReadOnlySpan reference, int basis, int weight, bool upsample, int maximumBasis)
- {
- ref short destinationBase = ref MemoryMarshal.GetReference(destination);
- ref short referenceBase = ref MemoryMarshal.GetReference(reference);
- int basisIncrement = upsample ? 2 : 1;
- int validCount = maximumBasis == int.MaxValue || basis >= maximumBasis
- ? maximumBasis == int.MaxValue ? destination.Length : 0
- : Math.Min(destination.Length, ((maximumBasis - 1 - basis) / basisIncrement) + 1);
- int index = 0;
-
- if (!upsample)
- {
- // High-bit-depth samples stay in signed 16-bit storage, but interpolation widens to Int32 before the Q5
- // weighted sum. The largest supported 12-bit sample therefore cannot overflow an intermediate lane.
- if (Vector512.IsHardwareAccelerated)
- {
- int oneVectorFromEnd = validCount - Vector512.Count;
- for (; index <= oneVectorFromEnd; index += Vector512.Count)
+ if (Vector128.IsHardwareAccelerated)
{
- Vector512 left = Vector512.LoadUnsafe(ref referenceBase, (nuint)(basis + index));
- Vector512 right = Vector512.LoadUnsafe(ref referenceBase, (nuint)(basis + index + 1));
- Interpolate(left, right, weight).StoreUnsafe(ref destinationBase, (nuint)index);
+ int oneVectorFromEnd = validCount - Vector128.Count;
+ for (; index <= oneVectorFromEnd; index += Vector128.Count)
+ {
+ Vector128 left = Vector128.LoadUnsafe(ref referenceBase, (nuint)(basis + index));
+ Vector128 right = Vector128.LoadUnsafe(ref referenceBase, (nuint)(basis + index + 1));
+ TOperator.Interpolate(left, right, weight).StoreUnsafe(ref destinationBase, (nuint)index);
+ }
}
}
-
- if (Vector256.IsHardwareAccelerated)
+ else if (Vector128.IsHardwareAccelerated)
{
- int oneVectorFromEnd = validCount - Vector256.Count;
- for (; index <= oneVectorFromEnd; index += Vector256.Count)
+ int oneVectorFromEnd = validCount - 4;
+ for (; index <= oneVectorFromEnd; index += 4)
{
- Vector256 left = Vector256.LoadUnsafe(ref referenceBase, (nuint)(basis + index));
- Vector256 right = Vector256.LoadUnsafe(ref referenceBase, (nuint)(basis + index + 1));
- Interpolate(left, right, weight).StoreUnsafe(ref destinationBase, (nuint)index);
+ Vector128 source = Vector128.LoadUnsafe(ref referenceBase, (nuint)(basis + (index * 2)));
+ Vector128 left = Vector128.ShuffleNative(source, EvenShortIndices);
+ Vector128 right = Vector128.ShuffleNative(source, OddShortIndices);
+ Vector128 prediction = TOperator.Interpolate(left, right, weight);
+ Unsafe.As(ref Unsafe.Add(ref destinationBase, index)) = prediction.AsUInt64().ToScalar();
}
}
if (Vector128.IsHardwareAccelerated)
{
- int oneVectorFromEnd = validCount - Vector128.Count;
- for (; index <= oneVectorFromEnd; index += Vector128.Count)
+ int oneVectorFromEnd = validCount - 4;
+ for (; index <= oneVectorFromEnd; index += 4)
{
- Vector128 left = Vector128.LoadUnsafe(ref referenceBase, (nuint)(basis + index));
- Vector128 right = Vector128.LoadUnsafe(ref referenceBase, (nuint)(basis + index + 1));
- Interpolate(left, right, weight).StoreUnsafe(ref destinationBase, (nuint)index);
+ int source = basis + (index * basisIncrement);
+ Vector128 left = Vector128.Create(
+ (int)Unsafe.Add(ref referenceBase, source),
+ Unsafe.Add(ref referenceBase, source + basisIncrement),
+ Unsafe.Add(ref referenceBase, source + (2 * basisIncrement)),
+ Unsafe.Add(ref referenceBase, source + (3 * basisIncrement)));
+
+ Vector128 right = Vector128.Create(
+ (int)Unsafe.Add(ref referenceBase, source + 1),
+ Unsafe.Add(ref referenceBase, source + basisIncrement + 1),
+ Unsafe.Add(ref referenceBase, source + (2 * basisIncrement) + 1),
+ Unsafe.Add(ref referenceBase, source + (3 * basisIncrement) + 1));
+
+ StoreFourShorts(TOperator.Interpolate(left, right, Vector128.Create(weight)), ref Unsafe.Add(ref destinationBase, index));
}
}
+
+ for (; index < validCount; index++)
+ {
+ int source = basis + (index * basisIncrement);
+ Unsafe.Add(ref destinationBase, index) = (short)(((Unsafe.Add(ref referenceBase, source) * (32 - weight)) + (Unsafe.Add(ref referenceBase, source + 1) * weight) + 16) >> 5);
+ }
+
+ if (validCount < destination.Length)
+ {
+ destination[validCount..].Fill(Unsafe.Add(ref referenceBase, maximumBasis));
+ }
}
- else if (Vector128.IsHardwareAccelerated)
+
+ ///
+ /// Interpolates the left-edge prefix of one 8-bit zone 2 row.
+ ///
+ /// The destination prefix.
+ /// The projected left reference.
+ /// The first Q6 left projection.
+ /// The Q8 derivative subtracted between columns.
+ /// Whether the left reference contains half-sample positions.
+ private static void InterpolateLeft(Span destination, ReadOnlySpan left, int projection, int derivative, bool upsample)
{
- int oneVectorFromEnd = validCount - 4;
- for (; index <= oneVectorFromEnd; index += 4)
+ ref byte destinationBase = ref MemoryMarshal.GetReference(destination);
+ ref byte leftBase = ref MemoryMarshal.GetReference(left);
+ int upsampleShift = upsample ? 1 : 0;
+ int fractionBits = 6 - upsampleShift;
+ int index = 0;
+
+ if (Vector128.IsHardwareAccelerated)
+ {
+ // Zone-two left references are not contiguous across output columns. Constructing the four source pairs
+ // directly avoids a temporary gather-index buffer and keeps the scalar continuation at the same offset.
+ int oneVectorFromEnd = destination.Length - 4;
+ for (; index <= oneVectorFromEnd; index += 4)
+ {
+ int projection0 = projection - (index * derivative);
+ int projection1 = projection0 - derivative;
+ int projection2 = projection1 - derivative;
+ int projection3 = projection2 - derivative;
+ int basis0 = projection0 >> fractionBits;
+ int basis1 = projection1 >> fractionBits;
+ int basis2 = projection2 >> fractionBits;
+ int basis3 = projection3 >> fractionBits;
+ Vector128 source0 = Vector128.Create((int)Unsafe.Add(ref leftBase, basis0), Unsafe.Add(ref leftBase, basis1), Unsafe.Add(ref leftBase, basis2), Unsafe.Add(ref leftBase, basis3));
+ Vector128 source1 = Vector128.Create((int)Unsafe.Add(ref leftBase, basis0 + 1), Unsafe.Add(ref leftBase, basis1 + 1), Unsafe.Add(ref leftBase, basis2 + 1), Unsafe.Add(ref leftBase, basis3 + 1));
+ Vector128 weights = Vector128.Create(
+ ((projection0 << upsampleShift) & 0x3F) >> 1,
+ ((projection1 << upsampleShift) & 0x3F) >> 1,
+ ((projection2 << upsampleShift) & 0x3F) >> 1,
+ ((projection3 << upsampleShift) & 0x3F) >> 1);
+
+ StoreFourBytes(TOperator.Interpolate(source0, source1, weights), ref Unsafe.Add(ref destinationBase, index));
+ }
+ }
+
+ for (; index < destination.Length; index++)
{
- Vector128 source = Vector128.LoadUnsafe(ref referenceBase, (nuint)(basis + (index * 2)));
- Vector128 left = Vector128.ShuffleNative(source, EvenShortIndices);
- Vector128 right = Vector128.ShuffleNative(source, OddShortIndices);
- Vector128 prediction = Interpolate(left, right, weight);
- Unsafe.As(ref Unsafe.Add(ref destinationBase, index)) = prediction.AsUInt64().GetElement(0);
+ int currentProjection = projection - (index * derivative);
+ int basis = currentProjection >> fractionBits;
+ int weight = ((currentProjection << upsampleShift) & 0x3F) >> 1;
+ Unsafe.Add(ref destinationBase, index) = (byte)(((Unsafe.Add(ref leftBase, basis) * (32 - weight)) + (Unsafe.Add(ref leftBase, basis + 1) * weight) + 16) >> 5);
}
}
- if (Vector128.IsHardwareAccelerated)
+ ///
+ /// Interpolates the left-edge prefix of one high-bit-depth zone 2 row.
+ ///
+ /// The destination prefix.
+ /// The projected left reference.
+ /// The first Q6 left projection.
+ /// The Q8 derivative subtracted between columns.
+ /// Whether the left reference contains half-sample positions.
+ private static void InterpolateLeft(Span destination, ReadOnlySpan left, int projection, int derivative, bool upsample)
{
- int oneVectorFromEnd = validCount - 4;
- for (; index <= oneVectorFromEnd; index += 4)
- {
- int source = basis + (index * basisIncrement);
- Vector128 left = Vector128.Create(
- (int)Unsafe.Add(ref referenceBase, source),
- Unsafe.Add(ref referenceBase, source + basisIncrement),
- Unsafe.Add(ref referenceBase, source + (2 * basisIncrement)),
- Unsafe.Add(ref referenceBase, source + (3 * basisIncrement)));
+ ref short destinationBase = ref MemoryMarshal.GetReference(destination);
+ ref short leftBase = ref MemoryMarshal.GetReference(left);
+ int upsampleShift = upsample ? 1 : 0;
+ int fractionBits = 6 - upsampleShift;
+ int index = 0;
- Vector128 right = Vector128.Create(
- (int)Unsafe.Add(ref referenceBase, source + 1),
- Unsafe.Add(ref referenceBase, source + basisIncrement + 1),
- Unsafe.Add(ref referenceBase, source + (2 * basisIncrement) + 1),
- Unsafe.Add(ref referenceBase, source + (3 * basisIncrement) + 1));
+ if (Vector128.IsHardwareAccelerated)
+ {
+ int oneVectorFromEnd = destination.Length - 4;
+ for (; index <= oneVectorFromEnd; index += 4)
+ {
+ int projection0 = projection - (index * derivative);
+ int projection1 = projection0 - derivative;
+ int projection2 = projection1 - derivative;
+ int projection3 = projection2 - derivative;
+ int basis0 = projection0 >> fractionBits;
+ int basis1 = projection1 >> fractionBits;
+ int basis2 = projection2 >> fractionBits;
+ int basis3 = projection3 >> fractionBits;
+ Vector128 source0 = Vector128.Create((int)Unsafe.Add(ref leftBase, basis0), Unsafe.Add(ref leftBase, basis1), Unsafe.Add(ref leftBase, basis2), Unsafe.Add(ref leftBase, basis3));
+ Vector128 source1 = Vector128.Create((int)Unsafe.Add(ref leftBase, basis0 + 1), Unsafe.Add(ref leftBase, basis1 + 1), Unsafe.Add(ref leftBase, basis2 + 1), Unsafe.Add(ref leftBase, basis3 + 1));
+ Vector128 weights = Vector128.Create(
+ ((projection0 << upsampleShift) & 0x3F) >> 1,
+ ((projection1 << upsampleShift) & 0x3F) >> 1,
+ ((projection2 << upsampleShift) & 0x3F) >> 1,
+ ((projection3 << upsampleShift) & 0x3F) >> 1);
+
+ StoreFourShorts(TOperator.Interpolate(source0, source1, weights), ref Unsafe.Add(ref destinationBase, index));
+ }
+ }
- StoreFourShorts(Interpolate(left, right, Vector128.Create(weight)), ref Unsafe.Add(ref destinationBase, index));
+ for (; index < destination.Length; index++)
+ {
+ int currentProjection = projection - (index * derivative);
+ int basis = currentProjection >> fractionBits;
+ int weight = ((currentProjection << upsampleShift) & 0x3F) >> 1;
+ Unsafe.Add(ref destinationBase, index) = (short)(((Unsafe.Add(ref leftBase, basis) * (32 - weight)) + (Unsafe.Add(ref leftBase, basis + 1) * weight) + 16) >> 5);
}
}
- for (; index < validCount; index++)
+ ///
+ /// Stores four widened predictions as packed 8-bit samples.
+ ///
+ /// The widened predictions.
+ /// The first destination sample.
+ [MethodImpl(MethodImplOptions.AggressiveInlining)]
+ private static void StoreFourBytes(Vector128 prediction, ref byte destination)
{
- int source = basis + (index * basisIncrement);
- Unsafe.Add(ref destinationBase, index) = (short)(((Unsafe.Add(ref referenceBase, source) * (32 - weight)) + (Unsafe.Add(ref referenceBase, source + 1) * weight) + 16) >> 5);
+ Vector128 narrowed16 = Vector128.Narrow(prediction.AsUInt32(), Vector128.Zero);
+ Vector128 narrowed8 = Vector128.Narrow(narrowed16, Vector128.Zero);
+ Unsafe.As(ref destination) = narrowed8.AsUInt32().GetElement(0);
}
- if (validCount < destination.Length)
+ ///
+ /// Stores four widened predictions as packed high-bit-depth samples.
+ ///
+ /// The widened predictions.
+ /// The first destination sample.
+ [MethodImpl(MethodImplOptions.AggressiveInlining)]
+ private static void StoreFourShorts(Vector128 prediction, ref short destination)
{
- destination[validCount..].Fill(Unsafe.Add(ref referenceBase, maximumBasis));
+ Vector128 narrowed = Vector128.Narrow(prediction, Vector128.Zero);
+ Unsafe.As(ref destination) = narrowed.AsUInt64().ToScalar();
}
}
///
- /// Interpolates the left-edge prefix of one 8-bit zone 2 row.
+ /// Implements directional intra prediction through one closed interpolation operator.
///
- /// The destination prefix.
- /// The projected left reference.
- /// The first Q6 left projection.
- /// The Q8 derivative subtracted between columns.
- /// Whether the left reference contains half-sample positions.
- private static void InterpolateLeft(Span destination, ReadOnlySpan left, int projection, int derivative, bool upsample)
+ /// The directional interpolation arithmetic.
+ private static partial class Predictor
+ where TOperator : struct, IDirectionalPredictionOperator
{
- ref byte destinationBase = ref MemoryMarshal.GetReference(destination);
- ref byte leftBase = ref MemoryMarshal.GetReference(left);
- int upsampleShift = upsample ? 1 : 0;
- int fractionBits = 6 - upsampleShift;
- int index = 0;
-
- if (Vector128.IsHardwareAccelerated)
+ ///
+ /// Gets the Q8 directional derivatives indexed by acute prediction angle.
+ ///
+ private static ReadOnlySpan DirectionalIntraDerivative =>
+ [
+
+ // Zero entries represent angles which AV1 never signals. Direct indexing avoids a search or division in
+ // each directional block while retaining the exact fixed-point projections from the normative table.
+ 0, 0, 0, 1023, 0, 0, 547, 0, 0, 372, 0, 0, 0, 0, 273, 0, 0, 215, 0, 0, 178, 0, 0,
+ 151, 0, 0, 132, 0, 0, 116, 0, 0, 102, 0, 0, 0, 90, 0, 0, 80, 0, 0, 71, 0, 0, 64, 0, 0,
+ 57, 0, 0, 51, 0, 0, 45, 0, 0, 0, 40, 0, 0, 35, 0, 0, 31, 0, 0, 27, 0, 0, 23, 0, 0,
+ 19, 0, 0, 15, 0, 0, 0, 0, 11, 0, 0, 7, 0, 0, 3, 0, 0,
+ ];
+
+ ///
+ /// Predicts an 8-bit directional block using the widest available SIMD path.
+ ///
+ /// The destination block origin.
+ /// The destination row stride in samples.
+ /// The predicted block dimensions.
+ /// The prepared top reference, including any required extension.
+ /// The prepared left reference, including any required extension.
+ /// Whether the top edge contains half-sample positions.
+ /// Whether the left edge contains half-sample positions.
+ /// The adjusted prediction angle.
+ /// The caller-owned block transposition workspace.
+ public static void Predict(Span destination, int destinationStride, Av1TransformSize transformSize, ReadOnlySpan above, ReadOnlySpan left, bool upsampleAbove, bool upsampleLeft, int angle, Span scratch)
{
- // Zone-two left references are not contiguous across output columns. Constructing the four source pairs
- // directly avoids a temporary gather-index buffer and keeps the scalar continuation at the same offset.
- int oneVectorFromEnd = destination.Length - 4;
- for (; index <= oneVectorFromEnd; index += 4)
- {
- int projection0 = projection - (index * derivative);
- int projection1 = projection0 - derivative;
- int projection2 = projection1 - derivative;
- int projection3 = projection2 - derivative;
- int basis0 = projection0 >> fractionBits;
- int basis1 = projection1 >> fractionBits;
- int basis2 = projection2 >> fractionBits;
- int basis3 = projection3 >> fractionBits;
- Vector128