From 930acdc82c0c66a268c6660a960469e55a7211c2 Mon Sep 17 00:00:00 2001 From: James Jackson-South Date: Fri, 4 Sep 2026 09:42:06 +1000 Subject: [PATCH] Implement AV1 lossless still-image coding --- HEIF_IMPLEMENTATION_PLAN.md | 8 +- .../Av1/OpenBitstreamUnit/ObuColorConfig.cs | 19 +- .../Av1EncoderModeDecisionWorkspace.cs | 4 +- .../Heif/Av1/Pipeline/Av1FrameEncoder.cs | 18 +- ...traSuperblockEncoder.ChromaModeDecision.cs | 20 ++- .../Av1IntraSuperblockEncoder.ModeDecision.cs | 52 ++++-- .../Av1/Pipeline/Av1TransformBlockEncoder.cs | 34 ++-- .../Av1ForwardQuantizer.Operator.cs | 170 ++++++++++++++++++ .../Quantizers/Av1ForwardQuantizer.cs | 49 ++++- .../Av1/Transform/Av1ForwardTransformer.cs | 127 +++++++++++++ .../Formats/Heif/HeifEncoderCore.cs | 15 +- .../Formats/Heif/Av1/Av1EncoderFrameTests.cs | 89 +++++++++ .../Heif/Av1/Av1ForwardQuantizerTests.cs | 35 +++- .../Heif/Av1/Av1ForwardTransformTests.cs | 91 ++++++++++ .../Heif/Av1/Av1TransformBlockEncoderTests.cs | 1 + .../Formats/Heif/HeifEncoderTests.cs | 62 ++++++- 16 files changed, 710 insertions(+), 84 deletions(-) diff --git a/HEIF_IMPLEMENTATION_PLAN.md b/HEIF_IMPLEMENTATION_PLAN.md index ba854d7280..834bc8ffe5 100644 --- a/HEIF_IMPLEMENTATION_PLAN.md +++ b/HEIF_IMPLEMENTATION_PLAN.md @@ -828,7 +828,7 @@ Encoder verification contract: - [~] Implement superblock and partition analysis for every permitted block size and partition. Efforts zero through eight deliberately split every in-frame node to 8x8 blocks. Effort nine performs recursive live rate-distortion selection at complete 8x8 and 16x16 nodes, while effort ten extends the same search to complete 32x32, 64x64, and 128x128 nodes. Candidate order matches current libaom: `PARTITION_NONE`, `PARTITION_SPLIT`, `PARTITION_HORZ`, `PARTITION_VERT`, the four asymmetric partitions, then `PARTITION_HORZ_4` and `PARTITION_VERT_4`; the two 1-to-4 partitions are excluded at 128x128 as required by current libaom. Invalid chroma geometries are excluded before evaluation. Each candidate saves and restores the exact partition, coefficient, transform, and palette neighbor edges in one aligned block-workspace owner; trials neither allocate nor copy probability state. Recursive split trials publish each selected child's decoded mode, transform, coefficient, and palette contexts before evaluating its next sibling. Coefficient contexts are published per retained transform rather than broadcasting the first transform over an entire partition leaf. Large luma and chroma leaves are evaluated as bounded-64, raster-ordered transform tiles in the existing aligned workspace, and each winning plane is copied to retained storage once. Production picture state retains the compact 8x8 mode allocation below effort nine and explicitly selects 4x4 allocation granularity when sub-8x8 partitions are enabled. Effort-dependent pruning remains. - [~] Implement intra mode search, palette, filter intra, chroma-from-luma, and intra-block copy decisions. Live luma search now covers all 13 zero-angle base modes and all six nonzero adjustments for each of the eight directional modes. Joint spatial chroma search covers the same 61 candidates, combines both chroma planes in one rate-distortion decision, and preserves the winning shared angle adjustment. Chroma-from-luma now searches the complete signed alpha alphabet from reconstructed luma and retains its joint U/V syntax. Filter-intra now searches all five predictors after ordinary luma modes. Palette entropy, retained state, production syntax, exhaustive luma and paired chroma palette selection, adaptive screen-content activation, and joint intra-block-copy mode selection are complete. - [ ] Implement inter mode search for bounded sequences, including reference selection and the decoder-supported inter tools. -- [~] Current-libaom `av1_quantize_fp_no_qmatrix` arithmetic is implemented as a closed generic forward-quantizer family with Vector512, Vector256, Vector128, and scalar paths, raster-order output, coded 64-point coefficient limits, and scan-order EOB selection. Transform search, coefficient optimization, and lossless behavior remain. +- [~] Current-libaom `av1_quantize_fp_no_qmatrix` arithmetic is implemented as a closed generic forward-quantizer family with Vector512, Vector256, Vector128, and scalar paths, raster-order output, coded 64-point coefficient limits, and scan-order EOB selection. High-bit-depth paths widen before multiplying instead of applying the eight-bit coefficient clamp. Lossless blocks use the AV1 4x4 Walsh-Hadamard transform, exact lossless quantization and dequantization, four-by-four-only transform syntax, and non-skipped residual coding. Transform search and coefficient optimization remain. - [~] Implement real rate-distortion selection and make quality and effort change work, size, and output quality. The complete luma and joint chroma candidate sets, including chroma-from-luma, filter-intra, palette, and intra-block copy, now perform live rate-distortion selection. Public quality mapping and effort tiers through exhaustive uniform luma mode/transform search are implemented. Effort nine adds exact recursive 8x8 and 16x16 partition rate-distortion selection, and effort ten extends it through 128x128; effort-dependent pruning and the remaining sequence searches remain. - [~] Frame effort now progressively expands the available current search: zero is DC-only, one adds every zero-angle spatial mode, two adds every legal directional adjustment, three refines the preliminary luma winner's transform type, four adds filter-intra and chroma-from-luma, and five adds adaptive palette and intra-block-copy analysis. Lower tiers do not signal unavailable sequence or frame tools, and tiers below five skip the whole-frame screen-content scan. Effort six enables `TX_MODE_SELECT` and compares the winning ordinary spatial or filter-intra luma mode as one 8x8 transform against four raster-ordered 4x4 transforms; each luma palette candidate owns that size comparison from effort six onward. Effort seven searches every legal 8x8 transform type inside every ordinary spatial candidate rather than refining only the preliminary winner. Effort eight also performs the 8x8-versus-four-4x4 comparison inside every ordinary spatial and filter-intra candidate, matching current libaom's per-candidate uniform-transform ownership. Effort nine additionally searches every legal partition at complete 8x8 and 16x16 nodes in current-libaom order, and effort ten extends that recursive search through 128x128. Prediction and residual construction run once per mode and are reused across its legal transform types. A 128x128 leaf evaluates four 64x64 luma transforms and as many as sixteen 32x32 transforms per 4:4:4 chroma plane, retaining sparse state at coefficient-area offsets. Residual emission follows AV1's bounded-region order, completing Y, U, and V for each 64x64 luma region before advancing. Every 4x4 transform searches all legal types with live coefficient contexts and reconstructed intra references. The search reuses the aligned block workspace, preserves only global improvements, and performs no per-block, per-partition, or per-transform rent. Non-skipped intra-block copy writes and costs the current-libaom unsplit variable-transform root; skipped intra-block copy emits no transform-partition symbol. Effort-dependent model and transform pruning remain. Decoder-visible production cases inspect the emitted restrictions and frame state and decode the produced streams, including real effort-nine streams selecting sub-8x8 and 8x16 rectangular blocks. The complete non-HEVC HEIF/AV1 namespace passes 9,077 of 9,077 through one foreground net11 Release VSTest run. The last independently built `aomdec`, from the then-current `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727` snapshot, accepts the previously generated effort-eight and effort-ten payloads as well as the existing palette and intra-block-copy payloads. The affected encoder, partition, and workspace surface passes 139 of 139 through one foreground net11 Release VSTest run. The net11 Release build and Roslynk compiler and analyzer passes report zero errors. - [~] Encoder rate accounting converts the entropy writer's live inverse cumulative distributions into current-libaom fixed-point symbol costs without allocating or duplicating probability state. Read-only luma-mode, directional-delta, filter-intra, chroma-mode, block-skip, transform-size, transform-block-skip, and complete transform-coefficient queries share the exact distributions mutated by the subsequent entropy write. Complete coefficient costing follows current libaom's optimized shape: it returns immediately for an empty transform, uses the EOB-specific base-range context, fuses magnitude, sign, base-range, and Golomb accounting into one reverse traversal, and combines repeated full base-range chunks instead of replaying each emitted symbol. Tile-lifetime level and context scratch is reused, the one-coefficient path neither clears nor initializes the forward-neighbor level map, and steady-state queries allocate nothing. Transform-size writing and costing share one subdivision-depth calculation, while shared closed symbol operations keep the writer and cost mappings for transform skip, transform type, and EOB syntax identical without forcing the estimator through the writer's slower two-pass coefficient traversal. The current-libaom fixed-point RD combiner preserves 64-bit distortion and rounds the weighted 1/512-bit rate at the required boundary. Its key-frame multiplier follows libaom's squared DC-quantizer formula and exact 10/12-bit normalization. Live final-block selection evaluates all 61 legal 8x8 luma candidates: the 13 zero-angle base modes in current-libaom order, followed by six nonzero adjustments for each directional mode. Joint chroma selection evaluates the equivalent 61 spatial candidates, combines U and V distortion plus coefficient rate, and charges one live chroma-mode and shared-angle symbol over the actual subsampled 4x4, 4x8, or 8x8 geometry. Chroma-from-luma subsamples the reconstructed luma block once into fixed-stride Q3 stack scratch, subtracts the rounded mean, evaluates all 33 signed alpha values independently for each plane with complete transform RD, and combines the cached plane results across all 1,088 valid joint pairs with one live sign cost and the conditional U/V magnitude costs. This is the allocation-free equivalent of current libaom's exhaustive 33-value path: it requires 66 evaluation transforms rather than transforming every joint pair, preserves DC-before-CfL-before-spatial tie order, and fixes the implicit chroma transform to DCT-DCT. Filter-intra follows ordinary luma candidates, searches all five predictors in syntax order, and evaluates every legal transform while reusing one prepared prediction and source residual per filter mode. Every candidate includes its live mode, angle, filter mode, alpha, and coefficient rate plus normalized pixel-domain distortion. Each prepared reference edge retains the common-corner prefix and twice the transform dimension required by directional prediction. A shared encoder/decoder availability calculation selects reconstructed top-right and bottom-left extensions according to tile, frame, superblock, and block reconstruction order; unavailable extensions repeat the nearest coded endpoint. Missing top or left edges retain current libaom's perpendicular-sample and bit-depth-midpoint rules. Directional prediction applies the AV1 three-degree adjustment step and reuses transform workspace for zone-three transposition before the transform overwrites it, keeping candidate evaluation allocation-free. The winning luma and chroma signed adjustments are retained in the packed final-block state consumed by the tile writer. The tile writer invokes these reusable workspace-backed selectors after mapping current neighbors and immediately before writing each block, so later decisions see reconstructed samples, coefficient contexts, and CDF updates from every preceding block. Block skip is read only after the callback has combined every coded plane. Luma and chroma candidate scratch is partitioned from the encoder's single aligned reusable block workspace; transform-size search uses that owner for four retained 4x4 transform states, local coefficient contexts, and the compact trial reconstruction needed to preserve the best result. No candidate path rents a buffer per block or per transform. Only a newly winning candidate is copied into retained frame storage. Production fixtures force every luma base predictor, both extreme adjustments in all three directional zones, available top-right and bottom-left extensions, high-bit-depth adjustment propagation, exact signed luma and chroma angle-rate terms, joint U/V decisions, packed chroma state, and 4:2:0, 4:2:2, and 4:4:4 transform geometry. The CfL fixtures derive target chroma from a pilot production encode's actual reconstructed luma through an independent scalar Q3 oracle and prove exact positive/negative alpha syntax plus zero-residual DCT-DCT reconstruction for all three subsampling geometries at 8, 10, and 12 bits. The stable fixed-DC traversal comparison uses neutral samples for which both the baseline and live search are contractually DC and skipped, instead of relying on textured content to happen to select the baseline mode. Luma palette selection now evaluates dominant-color and one-dimensional K-means candidates for every legal size, snaps near-cache colors with the reference threshold and tie order, removes duplicate snapped colors, extends boundary maps from active samples, and performs complete transform rate-distortion search. Ordinary DC and filter-intra candidates pay the palette-disabled symbol whenever screen-content syntax is enabled. The exact net11 Release rebuild reports 1,992 test-project warnings and zero errors, all 58 intra-superblock cases pass, all 8,935 AVIF cases pass, and all 230 HEIF cases pass. Remaining mode decision work includes transform-size coverage for filter-intra and palette, broader joint mode/transform refinement, and effort-dependent pruning. Non-empty intra blocks deliberately remain non-skipped, matching current libaom; later inter mode selection owns its distinct skip-transform RD decision. @@ -876,6 +876,8 @@ Encoder verification contract: - [x] Candidate distortion now follows the reference separation between immutable source residuals and reconstructed-pixel error. The forward-transform boundary accepts a read-only residual, so all transform types for one prediction reuse that block directly instead of copying it into transform scratch before every trial. The shared residual API now measures strided source-versus-reconstruction squared error with documented Vector512, Vector256, Vector128, and scalar traversal through ImageSharp's vector-count helpers, eliminating the former residual destination write and second reduction pass. The same path covers ordinary intra, filter intra, palette, chroma-from-luma, split transforms, and intra-block copy at 8, 10, and 12 bits without an allocation. Independent scalar, stride, tail, intrinsic-tier, and zero-allocation coverage passes with the 140-case focused encoder set; the complete non-HEVC HEIF/AV1 namespace passes 9,301 of 9,301. Representative palette, filter-intra, and effort-eight output hashes remain byte-identical, and current-main `aomdec` accepts every checked stream. +- [x] Lossless still-image coding now follows current libaom's qindex-zero path without introducing a per-block allocation or a second frame buffer. The forward 4x4 Walsh-Hadamard transform and its transpose into the entropy pipeline's row-major coefficient order are allocation-free at Vector512, Vector256, Vector128, and scalar tiers; lossless quantization reconstructs the original transform coefficient exactly. The mode decision fixes lossless transforms to DCT-DCT syntax and four-by-four blocks, disables transform skip for nonzero residuals, and excludes the fixed-eight-by-eight intra-block-copy search that cannot represent the required lossless transform grid. The frame coefficient owner reserves the exact worst-case 2,048 transform states needed by both 128x128 4:4:4 chroma planes without adding an allocation. The color configuration derives its plane count from the monochrome flag, so high-bit-depth direct-frame writers and readers cannot retain contradictory mutable state. Public eight-bit color and auxiliary-alpha round trips are pixel exact; direct 10-bit and 12-bit 4:4:4 frame round trips are exact at native-plane precision. The Release build completes with zero errors, Roslynk reports zero compiler errors, and the complete non-HEVC HEIF/AV1 namespace passes 9,081 of 9,081 through one foreground VSTest run. An independently built generic `aomdec` from current official libaom `main` at `d565eec60f084421fa34fc0534b760c6452b6a6c` accepts all 55 payloads regenerated by that run, including the new public color, auxiliary alpha, 10-bit, and 12-bit lossless streams. + ### 7. Write complete AVIF output - [~] The encoder-side AV1 codec configuration is now derived directly from the encoded sequence header and writes the fixed four-byte `av1C` record with empty `configOBUs`. The image payload retains the required sequence header, so the property introduces no sequence-header allocation, retention, or copy. Four production-header cases cover main, high, and professional profiles; 8-, 10-, and 12-bit precision; monochrome, 4:2:0, 4:2:2, and 4:4:4 sampling; exact fixed bytes; decoder reparsing; and header/property equivalence through direct net11 Release VSTest. Property-container emission and public AVIF activation remain open. @@ -894,8 +896,8 @@ Encoder verification contract: Encoder exit gate: -- [ ] Current-main libaom accepts every produced AV1 payload. -- [ ] Lossless output is exact at native-plane and final-pixel precision. +- [x] Current-main libaom accepts every currently produced AV1 payload. +- [~] Lossless output is exact for public eight-bit pixels and direct 8-, 10-, and 12-bit native planes; preservation of high-bit-depth public source precision remains part of the encoder contract work. - [ ] Lossy output demonstrates recorded quality and effort tradeoffs with absolute size, quality, timing, and allocation evidence. - [ ] 8, 10, and 12-bit monochrome, 4:2:0, 4:2:2, and 4:4:4 outputs pass. - [ ] Alpha, grids, metadata, color profiles, transforms, and bounded sequences pass. diff --git a/src/ImageSharp/Formats/Heif/Av1/OpenBitstreamUnit/ObuColorConfig.cs b/src/ImageSharp/Formats/Heif/Av1/OpenBitstreamUnit/ObuColorConfig.cs index 1b6f4b0970..4cac5ffdc8 100644 --- a/src/ImageSharp/Formats/Heif/Av1/OpenBitstreamUnit/ObuColorConfig.cs +++ b/src/ImageSharp/Formats/Heif/Av1/OpenBitstreamUnit/ObuColorConfig.cs @@ -8,11 +8,6 @@ namespace SixLabors.ImageSharp.Formats.Heif.Av1.OpenBitstreamUnit; /// internal sealed class ObuColorConfig { - /// - /// Stores whether the sequence uses a single monochrome plane. - /// - private bool isMonochrome; - /// /// Gets or sets a value indicating whether color-description syntax is present. /// @@ -21,23 +16,13 @@ internal sealed class ObuColorConfig /// /// Gets the number of color channels in this image. Can have the value 1 or 3. /// - public int PlaneCount { get; private set; } + public int PlaneCount => this.IsMonochrome ? 1 : Av1Constants.MaxPlanes; /// /// Gets or sets a value indicating whether the image has a single greyscale plane, will have /// color planes otherwise. /// - public bool IsMonochrome - { - get => this.isMonochrome; - set - { - // Plane count is derived from the monochrome flag throughout the decoder, so update - // both values atomically rather than allowing the two pieces of state to diverge. - this.PlaneCount = value ? 1 : Av1Constants.MaxPlanes; - this.isMonochrome = value; - } - } + public bool IsMonochrome { get; set; } /// /// Gets or sets the color-primary chromaticities. diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs index 2eedd1837c..e3079621f2 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs @@ -39,7 +39,8 @@ internal readonly ref struct Av1EncoderModeDecisionWorkspace /// /// The maximum number of transform states needed while evaluating both chroma planes of one 128x128 block. /// - public const int MaximumCandidateTransformBlockCount = 32; + public const int MaximumCandidateTransformBlockCount = + 2 * MaximumSampleCount / MinimumTransformSampleCount; /// /// The required workspace length in signed-integer storage elements. @@ -48,6 +49,7 @@ internal readonly ref struct Av1EncoderModeDecisionWorkspace private const int ReferenceBufferLength = (2 * Av1Constants.MaxTransformSize) + 1; private const int ReferenceBufferCount = 4; + private const int MinimumTransformSampleCount = 1 << (2 * Av1Constants.ModeInfoSizeLog2); private const int ReferenceStorageLength = ReferenceBufferCount * ReferenceBufferLength * sizeof(ushort) / sizeof(int); private const int CandidateSampleStorageOffset = ReferenceStorageLength; private const int CandidateSampleStorageLength = 2 * MaximumSampleCount * sizeof(ushort) / sizeof(int); diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1FrameEncoder.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1FrameEncoder.cs index 58769c65c2..01e7b5fcb5 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1FrameEncoder.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1FrameEncoder.cs @@ -119,7 +119,9 @@ internal static class Av1FrameEncoder ErrorResilientMode = true, RefreshFrameFlags = byte.MaxValue, DisableFrameEndUpdateCdf = true, - TransformMode = effort >= 6 ? Av1TransformMode.Select : Av1TransformMode.Largest, + TransformMode = qIndex == 0 + ? Av1TransformMode.Only4x4 + : effort >= 6 ? Av1TransformMode.Select : Av1TransformMode.Largest, ModeInfoColumnCount = modeInfoColumnCount, ModeInfoRowCount = modeInfoRowCount, TilesInfo = tiles, @@ -294,14 +296,17 @@ internal static class Av1FrameEncoder } frameHeader.AllowScreenContentTools = allowScreenContentTools; - frameHeader.AllowIntraBlockCopy = allowIntraBlockCopy; + + // The current IBC search is specialized for one 8x8 transform. Coded lossless requires reversible 4x4 + // transforms, so retain palette search but omit this candidate until it has a matching tiled implementation. + frameHeader.AllowIntraBlockCopy = !frameHeader.CodedLossless && allowIntraBlockCopy; using Av1EncoderPictureBuffer picture = new( configuration, sequenceHeader, frameHeader, image.Width, image.Height, - disallow4x4AllFrames: effort < 9); + disallow4x4AllFrames: !frameHeader.CodedLossless && effort < 9); using Av1EncoderCoefficientBuffer coefficients = new( configuration, @@ -358,14 +363,17 @@ internal static class Av1FrameEncoder } frameHeader.AllowScreenContentTools = allowScreenContentTools; - frameHeader.AllowIntraBlockCopy = allowIntraBlockCopy; + + // The current IBC search is specialized for one 8x8 transform. Coded lossless requires reversible 4x4 + // transforms, so retain palette search but omit this candidate until it has a matching tiled implementation. + frameHeader.AllowIntraBlockCopy = !frameHeader.CodedLossless && allowIntraBlockCopy; using Av1EncoderPictureBuffer picture = new( configuration, sequenceHeader, frameHeader, image.Width, image.Height, - disallow4x4AllFrames: effort < 9); + disallow4x4AllFrames: !frameHeader.CodedLossless && effort < 9); using Av1EncoderCoefficientBuffer coefficients = new( configuration, diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs index 13d32cb518..d7dbf10f9f 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ChromaModeDecision.cs @@ -805,10 +805,12 @@ internal static partial class Av1IntraSuperblockEncoder int transformHeight4x4 = transformSize.Get4x4HighCount(); int maximumUnitWidth = Math.Min(maximumUnitBlockSize.GetWidth(), blockWidth); int maximumUnitHeight = Math.Min(maximumUnitBlockSize.GetHeight(), blockHeight); - Av1TransformType transformType = Av1SymbolContextHelper.GetDefaultIntraTransformType( - predictionMode, - transformSize, - this.picture.Parent.FrameHeader.UseReducedTransformSet); + Av1TransformType transformType = this.picture.Parent.FrameHeader.CodedLossless + ? Av1TransformType.DctDct + : Av1SymbolContextHelper.GetDefaultIntraTransformType( + predictionMode, + transformSize, + this.picture.Parent.FrameHeader.UseReducedTransformSet); Av1ComponentType componentType = plane == Av1Plane.Y ? Av1ComponentType.Luminance @@ -1046,10 +1048,12 @@ internal static partial class Av1IntraSuperblockEncoder // Intra chroma derives one transform type from the shared UV prediction mode. The type is not // signaled independently for either chroma plane, so U and V must use the same legal fallback. - Av1TransformType transformType = Av1SymbolContextHelper.GetDefaultIntraTransformType( - predictionMode, - transformSize, - this.picture.Parent.FrameHeader.UseReducedTransformSet); + Av1TransformType transformType = this.picture.Parent.FrameHeader.CodedLossless + ? Av1TransformType.DctDct + : Av1SymbolContextHelper.GetDefaultIntraTransformType( + predictionMode, + transformSize, + this.picture.Parent.FrameHeader.UseReducedTransformSet); long distortion = TOperator.EncodeCandidate( this.blockWorkspace, diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs index aad34f945d..9692fef379 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs @@ -544,7 +544,10 @@ internal static partial class Av1IntraSuperblockEncoder { Av1BlockSize blockSize = modeInfo.Block.BlockSize; Av1PartitionType partitionType = modeInfo.Block.PartitionType; - Av1TransformSize maximumLumaTransformSize = blockSize.GetMaximumTransformSize(); + Av1TransformSize maximumLumaTransformSize = this.picture.Parent.FrameHeader.CodedLossless + ? Av1TransformSize.Size4x4 + : blockSize.GetMaximumTransformSize(); + int qIndex = this.quantization.QIndex[0]; modeInfo.Block = new Av1EncoderBlockModeInfo { @@ -619,8 +622,9 @@ internal static partial class Av1IntraSuperblockEncoder block.FilterIntraMode) : 0; - modeInfo.Block.Skip = lumaTransformEmpty && - Av1TileWriter.ShouldSkipCoefficients( + modeInfo.Block.Skip = + !this.picture.Parent.FrameHeader.CodedLossless && + lumaTransformEmpty && Av1TileWriter.ShouldSkipCoefficients( writer, Av1TileWriter.GetSkipContext(macroBlock), emptyTransformRate); @@ -660,9 +664,11 @@ internal static partial class Av1IntraSuperblockEncoder subsamplingX, subsamplingY); - Av1TransformSize chromaTransformSize = blockSize.GetMaxUvTransformSize( - colorConfig.SubSamplingX, - colorConfig.SubSamplingY); + Av1TransformSize chromaTransformSize = this.picture.Parent.FrameHeader.CodedLossless + ? Av1TransformSize.Size4x4 + : blockSize.GetMaxUvTransformSize( + colorConfig.SubSamplingX, + colorConfig.SubSamplingY); Av1BlockSize chromaBlockSize = blockSize.GetSubsampled( colorConfig.SubSamplingX, @@ -763,7 +769,8 @@ internal static partial class Av1IntraSuperblockEncoder Av1FilterIntraMode.AllFilterIntraModes); } - modeInfo.Block.Skip = Av1TileWriter.ShouldSkipCoefficients( + modeInfo.Block.Skip = !this.picture.Parent.FrameHeader.CodedLossless && + Av1TileWriter.ShouldSkipCoefficients( writer, Av1TileWriter.GetSkipContext(macroBlock), regularEmptyTransformRate); @@ -1336,7 +1343,11 @@ internal static partial class Av1IntraSuperblockEncoder out Av1TransformSize selectedTransformSize, out long selectedCost) { - Av1TransformSize transformSize = blockSize.GetMaximumTransformSize(); + bool codedLossless = this.picture.Parent.FrameHeader.CodedLossless; + Av1TransformSize transformSize = codedLossless + ? Av1TransformSize.Size4x4 + : blockSize.GetMaximumTransformSize(); + int blockWidth = blockSize.GetWidth(); int blockHeight = blockSize.GetHeight(); Av1EncoderModeDecisionWorkspace workspace = @@ -1530,8 +1541,9 @@ internal static partial class Av1IntraSuperblockEncoder // Transform type and transform size are separate search axes. Splitting their effort thresholds // gives callers a useful intermediate tier without changing the fast default path. - bool searchEveryTransformType = this.effort >= 7; - bool searchEveryTransformSize = this.effort >= 8 && + bool searchEveryTransformType = !codedLossless && this.effort >= 7; + bool searchEveryTransformSize = !codedLossless && + this.effort >= 8 && transformSize == Av1TransformSize.Size8x8; // Zero-angle modes precede groups of six nonzero adjustments for each directional mode. @@ -1571,12 +1583,14 @@ internal static partial class Av1IntraSuperblockEncoder // Transform types are visited in AV1 enumeration order. A strict cost comparison below keeps // the first legal type on ties, while lower efforts visit only the mode-derived default. - Av1TransformType firstTransformType = searchEveryTransformType + Av1TransformType firstTransformType = codedLossless ? Av1TransformType.DctDct - : Av1SymbolContextHelper.GetDefaultIntraTransformType( - mode, - transformSize, - useReducedTransformSet); + : searchEveryTransformType + ? Av1TransformType.DctDct + : Av1SymbolContextHelper.GetDefaultIntraTransformType( + mode, + transformSize, + useReducedTransformSet); Av1TransformType transformTypeLimit = searchEveryTransformType ? Av1TransformType.AllTransformTypes @@ -1679,7 +1693,7 @@ internal static partial class Av1IntraSuperblockEncoder // Midrange effort refines the preliminary mode only. Higher effort already searched every // mode-transform pair above, so repeating the winning mode would add no candidates. long bestTransformCost = bestCost; - if (this.effort >= 3 && !searchEveryTransformType) + if (!codedLossless && this.effort >= 3 && !searchEveryTransformType) { // The shared spans now contain the last mode visited above, so rebuild the preliminary // winner once before refining its transform types. @@ -1767,8 +1781,12 @@ internal static partial class Av1IntraSuperblockEncoder transformSize, this.bitDepth); + Av1TransformType filterTransformTypeLimit = codedLossless + ? (Av1TransformType)((int)Av1TransformType.DctDct + 1) + : Av1TransformType.AllTransformTypes; + for (Av1TransformType transformType = Av1TransformType.DctDct; - transformType < Av1TransformType.AllTransformTypes; + transformType < filterTransformTypeLimit; transformType++) { if (!transformType.IsExtendedSetUsed(transformSetType)) diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1TransformBlockEncoder.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1TransformBlockEncoder.cs index b2fecd255a..a61ec59951 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1TransformBlockEncoder.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1TransformBlockEncoder.cs @@ -223,10 +223,10 @@ internal static class Av1TransformBlockEncoder reconstruction, reconstructionStride, transformSize, - transformType, + state.TransformType, (int)plane, state.EndOfBlock, - false, + qIndex == 0, workspace.TransformWorkspace); } @@ -310,10 +310,10 @@ internal static class Av1TransformBlockEncoder reconstruction, width, transformSize, - Av1TransformType.DctDct, + state.TransformType, (int)plane, state.EndOfBlock, - false, + qIndex == 0, workspace.TransformWorkspace); } @@ -549,10 +549,10 @@ internal static class Av1TransformBlockEncoder MemoryMarshal.Cast(reconstruction), reconstructionStride, transformSize, - transformType, + state.TransformType, (int)plane, state.EndOfBlock, - false, + qIndex == 0, bitDepth, workspace.TransformWorkspace); } @@ -652,10 +652,10 @@ internal static class Av1TransformBlockEncoder signedReconstruction, width, transformSize, - Av1TransformType.DctDct, + state.TransformType, (int)plane, state.EndOfBlock, - false, + qIndex == 0, bitDepth, workspace.TransformWorkspace); } @@ -913,10 +913,10 @@ internal static class Av1TransformBlockEncoder reconstruction, reconstructionStride, transformSize, - transformType, + state.TransformType, (int)plane, state.EndOfBlock, - false, + qIndex == 0, workspace.TransformWorkspace); } } @@ -1001,10 +1001,10 @@ internal static class Av1TransformBlockEncoder MemoryMarshal.Cast(reconstruction), reconstructionStride, transformSize, - transformType, + state.TransformType, (int)plane, state.EndOfBlock, - false, + qIndex == 0, bitDepth, workspace.TransformWorkspace); } @@ -1061,6 +1061,16 @@ internal static class Av1TransformBlockEncoder Span quantized = quantizedCoefficients[..coefficientCount]; Span dequantized = workspace.DequantizedCoefficients[..coefficientCount]; + if (qIndex == 0) + { + // A coded-lossless frame fixes every transform block at 4x4 and uses the reversible transform. Its + // quantizer removes only the transform's fixed scale, leaving reconstruction coefficients unchanged. + Av1ForwardTransformer.TransformLossless4x4(residual, transformed, (uint)transformSize.GetWidth()); + state.EndOfBlock = Av1ForwardQuantizer.QuantizeLossless(transformed, quantized, dequantized, bitDepth); + state.TransformType = Av1TransformType.DctDct; + return; + } + // The forward transform reads the prepared residual without changing it, so every type candidate can // reuse one source-minus-prediction block. Separate coefficient spans preserve each later representation. Av1ForwardTransformer.Transform2d( diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Quantizers/Av1ForwardQuantizer.Operator.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Quantizers/Av1ForwardQuantizer.Operator.cs index 114b69d60e..98fc6e96cd 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Quantizers/Av1ForwardQuantizer.Operator.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Quantizers/Av1ForwardQuantizer.Operator.cs @@ -202,4 +202,174 @@ internal static partial class Av1ForwardQuantizer return quantized; } } + + /// + /// Implements fast no-matrix quantization without truncating high-bit-depth transform magnitudes. + /// + /// + /// The reciprocal product is widened to 64 bits because 10-bit and 12-bit transforms can exceed the signed + /// 16-bit range. Each SIMD overload preserves the scalar fixed-point operation order in every lane. + /// + internal readonly struct HighBitDepthFastQuantizationOperator : IForwardQuantizationOperator + { + /// + public static Vector128 Quantize( + Vector128 coefficients, + Vector128 rounding, + Vector128 quantizer, + Vector128 dequantizer, + int logScale, + out Vector128 dequantizedCoefficients) + { + Vector128 coefficientSign = Vector128.ShiftRightArithmetic(coefficients, 31); + Vector128 magnitude = Vector128.Abs(coefficients); + Vector128 thresholdMask = ~Vector128.GreaterThan(dequantizer, magnitude << (1 + logScale)); + Vector128 rounded = magnitude + rounding; + + // Widen before multiplying so transform magnitudes above 32,767 retain their full precision. + Vector128 lowerProduct = Vector128.WidenLower(rounded) * Vector128.WidenLower(quantizer); + Vector128 upperProduct = Vector128.WidenUpper(rounded) * Vector128.WidenUpper(quantizer); + Vector128 quantizedMagnitude = + Vector128.Narrow(lowerProduct >> (16 - logScale), upperProduct >> (16 - logScale)) & thresholdMask; + + Vector128 quantized = (quantizedMagnitude ^ coefficientSign) - coefficientSign; + Vector128 dequantizedMagnitude = (quantizedMagnitude * dequantizer) >> logScale; + dequantizedCoefficients = (dequantizedMagnitude ^ coefficientSign) - coefficientSign; + return quantized; + } + + /// + public static Vector256 Quantize( + Vector256 coefficients, + Vector256 rounding, + Vector256 quantizer, + Vector256 dequantizer, + int logScale, + out Vector256 dequantizedCoefficients) + { + Vector256 coefficientSign = Vector256.ShiftRightArithmetic(coefficients, 31); + Vector256 magnitude = Vector256.Abs(coefficients); + Vector256 thresholdMask = ~Vector256.GreaterThan(dequantizer, magnitude << (1 + logScale)); + Vector256 rounded = magnitude + rounding; + + // Widen before multiplying so transform magnitudes above 32,767 retain their full precision. + Vector256 lowerProduct = Vector256.WidenLower(rounded) * Vector256.WidenLower(quantizer); + Vector256 upperProduct = Vector256.WidenUpper(rounded) * Vector256.WidenUpper(quantizer); + Vector256 quantizedMagnitude = + Vector256.Narrow(lowerProduct >> (16 - logScale), upperProduct >> (16 - logScale)) & thresholdMask; + + Vector256 quantized = (quantizedMagnitude ^ coefficientSign) - coefficientSign; + Vector256 dequantizedMagnitude = (quantizedMagnitude * dequantizer) >> logScale; + dequantizedCoefficients = (dequantizedMagnitude ^ coefficientSign) - coefficientSign; + return quantized; + } + + /// + public static Vector512 Quantize( + Vector512 coefficients, + Vector512 rounding, + Vector512 quantizer, + Vector512 dequantizer, + int logScale, + out Vector512 dequantizedCoefficients) + { + Vector512 coefficientSign = Vector512.ShiftRightArithmetic(coefficients, 31); + Vector512 magnitude = Vector512.Abs(coefficients); + Vector512 thresholdMask = ~Vector512.GreaterThan(dequantizer, magnitude << (1 + logScale)); + Vector512 rounded = magnitude + rounding; + + // Widen before multiplying so transform magnitudes above 32,767 retain their full precision. + Vector512 lowerProduct = Vector512.WidenLower(rounded) * Vector512.WidenLower(quantizer); + Vector512 upperProduct = Vector512.WidenUpper(rounded) * Vector512.WidenUpper(quantizer); + Vector512 quantizedMagnitude = + Vector512.Narrow(lowerProduct >> (16 - logScale), upperProduct >> (16 - logScale)) & thresholdMask; + + Vector512 quantized = (quantizedMagnitude ^ coefficientSign) - coefficientSign; + Vector512 dequantizedMagnitude = (quantizedMagnitude * dequantizer) >> logScale; + dequantizedCoefficients = (dequantizedMagnitude ^ coefficientSign) - coefficientSign; + return quantized; + } + + /// + public static int Quantize( + int coefficient, + int rounding, + int quantizer, + int dequantizer, + int logScale, + out int dequantizedCoefficient) + { + int coefficientSign = coefficient >> 31; + int magnitude = (coefficient ^ coefficientSign) - coefficientSign; + int quantizedMagnitude = 0; + + if (((long)magnitude << (1 + logScale)) >= dequantizer) + { + quantizedMagnitude = (int)((((long)magnitude + rounding) * quantizer) >> (16 - logScale)); + } + + int quantized = (quantizedMagnitude ^ coefficientSign) - coefficientSign; + int dequantizedMagnitude = (quantizedMagnitude * dequantizer) >> logScale; + dequantizedCoefficient = (dequantizedMagnitude ^ coefficientSign) - coefficientSign; + return quantized; + } + } + + /// + /// Removes the reversible transform's fixed scale for entropy coding while retaining exact reconstruction coefficients. + /// + internal readonly struct LosslessQuantizationOperator : IForwardQuantizationOperator + { + /// + public static Vector128 Quantize( + Vector128 coefficients, + Vector128 rounding, + Vector128 quantizer, + Vector128 dequantizer, + int logScale, + out Vector128 dequantizedCoefficients) + { + dequantizedCoefficients = coefficients; + return coefficients >> 2; + } + + /// + public static Vector256 Quantize( + Vector256 coefficients, + Vector256 rounding, + Vector256 quantizer, + Vector256 dequantizer, + int logScale, + out Vector256 dequantizedCoefficients) + { + dequantizedCoefficients = coefficients; + return coefficients >> 2; + } + + /// + public static Vector512 Quantize( + Vector512 coefficients, + Vector512 rounding, + Vector512 quantizer, + Vector512 dequantizer, + int logScale, + out Vector512 dequantizedCoefficients) + { + dequantizedCoefficients = coefficients; + return coefficients >> 2; + } + + /// + public static int Quantize( + int coefficient, + int rounding, + int quantizer, + int dequantizer, + int logScale, + out int dequantizedCoefficient) + { + dequantizedCoefficient = coefficient; + return coefficient >> 2; + } + } } diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Quantizers/Av1ForwardQuantizer.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Quantizers/Av1ForwardQuantizer.cs index 6395598892..76341f56c6 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Quantizers/Av1ForwardQuantizer.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Quantizers/Av1ForwardQuantizer.cs @@ -36,21 +36,56 @@ internal static partial class Av1ForwardQuantizer int dcDeltaQ, int acDeltaQ, Av1BitDepth bitDepth) - => QuantizeLossy( + => bitDepth == Av1BitDepth.EightBit + ? Quantize( + coefficients, + quantizedCoefficients, + dequantizedCoefficients, + transformSize, + transformType, + qIndex, + dcDeltaQ, + acDeltaQ, + bitDepth) + : Quantize( + coefficients, + quantizedCoefficients, + dequantizedCoefficients, + transformSize, + transformType, + qIndex, + dcDeltaQ, + acDeltaQ, + bitDepth); + + /// + /// Quantizes one reversible four-by-four transform without changing its reconstruction coefficients. + /// + /// The raster-order reversible-transform coefficients. + /// The raster-order entropy-coding coefficients. + /// The raster-order reconstruction coefficients. + /// The coded sample bit depth. + /// The one-based end position in coefficient scan order. + public static ushort QuantizeLossless( + ReadOnlySpan coefficients, + Span quantizedCoefficients, + Span dequantizedCoefficients, + Av1BitDepth bitDepth) + => Quantize( coefficients, quantizedCoefficients, dequantizedCoefficients, - transformSize, - transformType, - qIndex, - dcDeltaQ, - acDeltaQ, + Av1TransformSize.Size4x4, + Av1TransformType.DctDct, + 0, + 0, + 0, bitDepth); /// /// Applies a closed generic quantization operator across the widest available hardware widths. /// - internal static ushort QuantizeLossy( + private static ushort Quantize( ReadOnlySpan coefficients, Span quantizedCoefficients, Span dequantizedCoefficients, diff --git a/src/ImageSharp/Formats/Heif/Av1/Transform/Av1ForwardTransformer.cs b/src/ImageSharp/Formats/Heif/Av1/Transform/Av1ForwardTransformer.cs index 4b58919818..14671c8cba 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Transform/Av1ForwardTransformer.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Transform/Av1ForwardTransformer.cs @@ -45,6 +45,133 @@ internal static partial class Av1ForwardTransformer DispatchColumn(input, coefficients, stride, bitDepth, ref config, workspace); } + /// + /// Applies the reversible four-by-four transform required by coded-lossless AV1 blocks. + /// + /// The spatial residual samples. + /// The destination transform coefficients. + /// The number of input samples between rows. + public static void TransformLossless4x4(ReadOnlySpan input, Span coefficients, uint stride) + { + int inputStride = (int)stride; + + if (Vector128.IsHardwareAccelerated) + { + Vector128 row0 = Vector128.Create((int)input[0], input[1], input[2], input[3]); + Vector128 row1 = Vector128.Create( + input[inputStride], + input[inputStride + 1], + input[inputStride + 2], + input[inputStride + 3]); + + Vector128 row2 = Vector128.Create( + input[2 * inputStride], + input[(2 * inputStride) + 1], + input[(2 * inputStride) + 2], + input[(2 * inputStride) + 3]); + + Vector128 row3 = Vector128.Create( + input[3 * inputStride], + input[(3 * inputStride) + 1], + input[(3 * inputStride) + 2], + input[(3 * inputStride) + 3]); + + // Each lane initially holds one column. The first stage transforms those columns in parallel, then the + // transpose makes each transformed column a row so the same lane-wise network can process the other axis. + TransformLosslessStage(ref row0, ref row1, ref row2, ref row3); + Av1Transform2dOperations.Transpose(ref row0, ref row1, ref row2, ref row3); + TransformLosslessStage(ref row0, ref row1, ref row2, ref row3); + + // The entropy pipeline stores transform positions in row-major order, while the reversible reference + // walk leaves the two frequency axes exchanged. Normalize that boundary before scan-order traversal. + Av1Transform2dOperations.Transpose(ref row0, ref row1, ref row2, ref row3); + + ref int coefficientBase = ref MemoryMarshal.GetReference(coefficients); + (row0 * 4).StoreUnsafe(ref coefficientBase); + (row1 * 4).StoreUnsafe(ref coefficientBase, 4); + (row2 * 4).StoreUnsafe(ref coefficientBase, 8); + (row3 * 4).StoreUnsafe(ref coefficientBase, 12); + return; + } + + // The first pass writes transposed columns into the destination, matching the layout consumed in-place by + // the second pass. This keeps the scalar fallback allocation-free without a temporary matrix. + for (int column = 0; column < 4; column++) + { + int a = input[column]; + int b = input[inputStride + column]; + int c = input[(2 * inputStride) + column]; + int d = input[(3 * inputStride) + column]; + + a += b; + d -= c; + int e = (a - d) >> 1; + b = e - b; + c = e - c; + a -= c; + d += b; + + int offset = column * 4; + coefficients[offset] = a; + coefficients[offset + 1] = c; + coefficients[offset + 2] = d; + coefficients[offset + 3] = b; + } + + for (int column = 0; column < 4; column++) + { + int a = coefficients[column]; + int b = coefficients[4 + column]; + int c = coefficients[8 + column]; + int d = coefficients[12 + column]; + + a += b; + d -= c; + int e = (a - d) >> 1; + b = e - b; + c = e - c; + a -= c; + d += b; + + coefficients[column] = a * 4; + coefficients[4 + column] = c * 4; + coefficients[8 + column] = d * 4; + coefficients[12 + column] = b * 4; + } + + // Normalize the scalar reference walk to the row-major coefficient contract used by entropy coding. + (coefficients[1], coefficients[4]) = (coefficients[4], coefficients[1]); + (coefficients[2], coefficients[8]) = (coefficients[8], coefficients[2]); + (coefficients[3], coefficients[12]) = (coefficients[12], coefficients[3]); + (coefficients[6], coefficients[9]) = (coefficients[9], coefficients[6]); + (coefficients[7], coefficients[13]) = (coefficients[13], coefficients[7]); + (coefficients[11], coefficients[14]) = (coefficients[14], coefficients[11]); + } + + /// + /// Applies one axis of the reversible four-point transform to four independent SIMD lanes. + /// + [MethodImpl(MethodImplOptions.AggressiveInlining)] + private static void TransformLosslessStage( + ref Vector128 row0, + ref Vector128 row1, + ref Vector128 row2, + ref Vector128 row3) + { + Vector128 a = row0 + row1; + Vector128 d = row3 - row2; + Vector128 e = (a - d) >> 1; + Vector128 b = e - row1; + Vector128 c = e - row2; + a -= c; + d += b; + + row0 = a; + row1 = c; + row2 = d; + row3 = b; + } + /// /// Selects the concrete column operator for a transform block. /// diff --git a/src/ImageSharp/Formats/Heif/HeifEncoderCore.cs b/src/ImageSharp/Formats/Heif/HeifEncoderCore.cs index 1f21a76fa8..d7f267c449 100644 --- a/src/ImageSharp/Formats/Heif/HeifEncoderCore.cs +++ b/src/ImageSharp/Formats/Heif/HeifEncoderCore.cs @@ -786,11 +786,6 @@ internal sealed class HeifEncoderCore CancellationToken cancellationToken) where TPixel : unmanaged, IPixel { - if (this.encoder.Lossless) - { - throw new NotSupportedException("Lossless AV1 encoding is not implemented."); - } - if (image.Frames.Count != 1) { throw new NotSupportedException("AV1 image-sequence encoding is not implemented."); @@ -806,8 +801,12 @@ internal sealed class HeifEncoderCore _ => throw new NotSupportedException($"HEIF bit depth '{bitDepth}' is not supported.") }; + HeifChromaSubsampling defaultChromaSubsampling = this.encoder.Lossless + ? HeifChromaSubsampling.Yuv444 + : HeifChromaSubsampling.Yuv420; + HeifChromaSubsampling chromaSubsampling = this.encoder.ChromaSubsampling ?? - (metadata.IsMonochrome ? HeifChromaSubsampling.Monochrome : HeifChromaSubsampling.Yuv420); + (metadata.IsMonochrome ? HeifChromaSubsampling.Monochrome : defaultChromaSubsampling); (bool isMonochrome, bool subsamplingX, bool subsamplingY) = chromaSubsampling switch { @@ -871,7 +870,7 @@ internal sealed class HeifEncoderCore }; int quality = this.encoder.Quality ?? 75; - int qIndex = GetAv1QuantizerIndex(quality); + int qIndex = this.encoder.Lossless ? 0 : GetAv1QuantizerIndex(quality); cancellationToken.ThrowIfCancellationRequested(); ObuSequenceHeader colorHeader = Av1FrameEncoder.Encode( this.configuration, @@ -910,7 +909,7 @@ internal sealed class HeifEncoderCore }; int alphaQuality = this.encoder.AlphaQuality ?? quality; - int alphaQIndex = GetAv1QuantizerIndex(alphaQuality); + int alphaQIndex = this.encoder.Lossless ? 0 : GetAv1QuantizerIndex(alphaQuality); cancellationToken.ThrowIfCancellationRequested(); long alphaOffset = stream.Length; ObuSequenceHeader alphaHeader = Av1FrameEncoder.EncodeAlpha( diff --git a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs index c66b6641bc..79e49ca68f 100644 --- a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs +++ b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs @@ -153,6 +153,95 @@ public class Av1EncoderFrameTests } } + [Theory] + [InlineData(TenBit)] + [InlineData(TwelveBit)] + public void LosslessHighBitDepthEncodingPreservesNativePlanes(int bitDepthValue) + { + const int width = 8; + const int height = 8; + Av1BitDepth bitDepth = (Av1BitDepth)bitDepthValue; + using Image source = new(width, height); + for (int row = 0; row < height; row++) + { + Span pixels = source.Frames.RootFrame.PixelBuffer.DangerousGetRowSpan(row); + for (int column = 0; column < width; column++) + { + pixels[column] = new Rgb48( + (ushort)((column * 7001) + (row * 997)), + (ushort)((row * 6007) + (column * 1231)), + (ushort)((column * 4001) + (row * 3001))); + } + } + + ObuColorConfig colorConfig = new() + { + IsColorDescriptionPresent = true, + ColorPrimaries = ObuColorPrimaries.Bt709, + TransferCharacteristics = ObuTransferCharacteristics.Srgb, + MatrixCoefficients = ObuMatrixCoefficients.Identity, + ColorRange = true, + BitDepth = bitDepth + }; + + using Av1EncoderFrameBuffer expected = new( + Configuration.Default, + width, + height, + bitDepth.GetBitCount(), + Av1ColorFormat.Yuv444, + 1, + 1); + + Av1FrameEncoder.PrepareSource( + Configuration.Default, + source.Frames.RootFrame, + expected.Frame, + colorConfig); + + using MemoryStream stream = new(); + Av1FrameEncoder.Encode( + Configuration.Default, + source.Frames.RootFrame, + stream, + colorConfig, + qIndex: 0, + effort: 0); + + byte[] payload = stream.ToArray(); + string outputDirectory = Path.Combine( + TestEnvironment.ActualOutputDirectoryFullPath, + "Formats", + "Heif", + "Av1"); + + Directory.CreateDirectory(outputDirectory); + File.WriteAllBytes( + Path.Combine(outputDirectory, $"encoder-frame-8x8-{bitDepth.GetBitCount()}b-444-lossless.obu"), + payload); + + using Av1Decoder decoder = new(Configuration.Default); + using Av1FrameBuffer actual = decoder.DecodeFrameBuffer(payload, null, null, out _); + ObuFrameHeader frameHeader = Assert.IsType(decoder.FrameHeader); + + Assert.Equal(width, actual.Width); + Assert.Equal(height, actual.Height); + Assert.True(frameHeader.CodedLossless); + Assert.True(frameHeader.AllLossless); + Assert.Equal(Av1TransformMode.Only4x4, frameHeader.TransformMode); + + foreach (Av1Plane plane in new[] { Av1Plane.Y, Av1Plane.U, Av1Plane.V }) + { + Buffer2DRegion expectedPlane = expected.Frame.View.GetPlane(plane); + for (int row = 0; row < height; row++) + { + Assert.Equal( + expectedPlane.DangerousGetRowSpan(row)[..width].ToArray(), + actual.GetHighBitDepthRowSpan(plane, row, 0, 0).ToArray()); + } + } + } + [Fact] public void EncodeEffortNineSelectsSubEightPartition() { diff --git a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1ForwardQuantizerTests.cs b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1ForwardQuantizerTests.cs index ae488bb0b8..1fcce27ff2 100644 --- a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1ForwardQuantizerTests.cs +++ b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1ForwardQuantizerTests.cs @@ -68,6 +68,34 @@ public class Av1ForwardQuantizerTests Assert.Equal(0, GC.GetAllocatedBytesForCurrentThread() - before); } + /// + /// Verifies that lossless quantization removes only the reversible transform scale. + /// + [Fact] + public void LosslessQuantizerRetainsExactReconstructionCoefficients() + { + int[] coefficients = + [ + 4, -8, 12, -16, + 20, -24, 28, -32, + 36, -40, 44, -48, + 52, -56, 60, -64 + ]; + + int[] quantized = new int[coefficients.Length]; + int[] dequantized = new int[coefficients.Length]; + + ushort endOfBlock = Av1ForwardQuantizer.QuantizeLossless( + coefficients, + quantized, + dequantized, + Av1BitDepth.TwelveBit); + + Assert.Equal((ushort)16, endOfBlock); + Assert.Equal(coefficients.Select(x => x / 4), quantized); + Assert.Equal(coefficients, dequantized); + } + /// /// Exercises each transform-scale category, coded 64-point layout, quantizer range, and sample precision. /// @@ -188,7 +216,12 @@ public class Av1ForwardQuantizerTests if ((magnitude << (1 + logScale)) >= dequantizer) { - magnitude = Math.Clamp(magnitude + rounding, short.MinValue, short.MaxValue); + magnitude += rounding; + if (bitDepth == Av1BitDepth.EightBit) + { + magnitude = Math.Min(magnitude, short.MaxValue); + } + quantizedMagnitude = (int)((magnitude * quantizer) >> (16 - logScale)); } diff --git a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1ForwardTransformTests.cs b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1ForwardTransformTests.cs index f5c9a74472..7cd8097fd7 100644 --- a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1ForwardTransformTests.cs +++ b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1ForwardTransformTests.cs @@ -61,6 +61,13 @@ public class Av1ForwardTransformTests public void TwoDimensionalPipelineMatchesReference() => FeatureTestRunner.RunWithHwIntrinsicsFeature(AssertTwoDimensionalPipeline, TransformConfigurations); + /// + /// Verifies the reversible transform against the independent scalar operation order at every hardware tier. + /// + [Fact] + public void LosslessTransformMatchesLibaomReferenceAcrossHardwareWidths() + => FeatureTestRunner.RunWithHwIntrinsicsFeature(AssertLosslessTransform, TransformConfigurations); + /// /// Verifies that the complete transform dispatcher reuses caller-owned workspace. /// @@ -83,6 +90,90 @@ public class Av1ForwardTransformTests Assert.Equal(0, GC.GetAllocatedBytesForCurrentThread() - before); } + private static void AssertLosslessTransform() + { + const int stride = 7; + short[] input = new short[stride * 4]; + for (int row = 0; row < 4; row++) + { + for (int column = 0; column < 4; column++) + { + input[(row * stride) + column] = (short)((row * 1003) - (column * 499) + (row * column * 71) - 1024); + } + } + + int[] expected = new int[16]; + int[] actual = new int[16]; + TransformLosslessReference(input, expected, stride); + Av1ForwardTransformer.TransformLossless4x4(input, actual, stride); + + Assert.Equal(expected, actual); + + long before = GC.GetAllocatedBytesForCurrentThread(); + for (int iteration = 0; iteration < 32; iteration++) + { + Av1ForwardTransformer.TransformLossless4x4(input, actual, stride); + } + + Assert.Equal(0, GC.GetAllocatedBytesForCurrentThread() - before); + } + + private static void TransformLosslessReference(ReadOnlySpan input, Span output, int stride) + { + // The first pass traverses columns and writes their four transformed values contiguously. The second + // pass consumes that transposed layout in place, matching the normative reversible operation order. + for (int column = 0; column < 4; column++) + { + int a = input[column]; + int b = input[stride + column]; + int c = input[(2 * stride) + column]; + int d = input[(3 * stride) + column]; + + a += b; + d -= c; + int e = (a - d) >> 1; + b = e - b; + c = e - c; + a -= c; + d += b; + + int offset = column * 4; + output[offset] = a; + output[offset + 1] = c; + output[offset + 2] = d; + output[offset + 3] = b; + } + + for (int column = 0; column < 4; column++) + { + int a = output[column]; + int b = output[4 + column]; + int c = output[8 + column]; + int d = output[12 + column]; + + a += b; + d -= c; + int e = (a - d) >> 1; + b = e - b; + c = e - c; + a -= c; + d += b; + + output[column] = a * 4; + output[4 + column] = c * 4; + output[8 + column] = d * 4; + output[12 + column] = b * 4; + } + + // Encoder coefficient storage is row-major, so exchange the reference walk's final frequency axes. + (output[1], output[4]) = (output[4], output[1]); + (output[2], output[8]) = (output[8], output[2]); + (output[3], output[12]) = (output[12], output[3]); + (output[6], output[9]) = (output[9], output[6]); + (output[7], output[13]) = (output[13], output[7]); + (output[11], output[14]) = (output[14], output[11]); + } + /// /// Exercises every DCT, ADST, and identity stage network using both Int16 and Int32 lane arithmetic. /// diff --git a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs index 62743aa865..4fe759a339 100644 --- a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs +++ b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs @@ -626,6 +626,7 @@ public class Av1TransformBlockEncoderTests Assert.Equal(Av1EncoderModeDecisionWorkspace.MaximumTransformSampleCount, modeWorkspace.Prediction.Length); Assert.Equal(Av1EncoderModeDecisionWorkspace.MaximumTransformSampleCount, modeWorkspace.Residual.Length); Assert.Equal(Av1EncoderModeDecisionWorkspace.MaximumCandidateTransformBlockCount, modeWorkspace.CandidateTransformBlocks.Length); + Assert.Equal(2048, modeWorkspace.CandidateTransformBlocks.Length); // CfL is unavailable above 32x32, so its scratch remains fixed while larger partitions are enabled. Assert.Equal(Av1ChromaFromLumaContext.BufferLength, modeWorkspace.ChromaFromLumaSamples.Length); diff --git a/tests/ImageSharp.Tests/Formats/Heif/HeifEncoderTests.cs b/tests/ImageSharp.Tests/Formats/Heif/HeifEncoderTests.cs index 8773b36789..404ea5f98d 100644 --- a/tests/ImageSharp.Tests/Formats/Heif/HeifEncoderTests.cs +++ b/tests/ImageSharp.Tests/Formats/Heif/HeifEncoderTests.cs @@ -8,6 +8,7 @@ using SixLabors.ImageSharp.Formats; using SixLabors.ImageSharp.Formats.Heif; using SixLabors.ImageSharp.Formats.Heif.Av1; using SixLabors.ImageSharp.Formats.Heif.Av1.OpenBitstreamUnit; +using SixLabors.ImageSharp.Formats.Heif.Av1.Transform; using SixLabors.ImageSharp.Memory; using SixLabors.ImageSharp.Metadata; using SixLabors.ImageSharp.Metadata.Profiles.Cicp; @@ -181,18 +182,69 @@ public class HeifEncoderTests } [Fact] - public void Av1RejectsLosslessEncodingBeforeWritingOutput() + public void Av1LosslessRoundTripPreservesColorAndAlpha() { - using Image image = new(1, 1); + const int width = 8; + const int height = 8; + using Image image = new(width, height); + for (int row = 0; row < height; row++) + { + Span pixels = image.Frames.RootFrame.PixelBuffer.DangerousGetRowSpan(row); + for (int column = 0; column < width; column++) + { + pixels[column] = new Rgba32( + (byte)((column * 31) + row), + (byte)((row * 29) + column), + (byte)((column * 17) + (row * 11)), + (byte)((column * 23) + (row * 7))); + } + } + + // Identity-matrix 4:4:4 maps the packed RGB channels directly onto AV1 planes, so codec losslessness + // can be asserted against the original pixels without a separate color-conversion tolerance. + image.Metadata.CicpProfile = new CicpProfile(1, 13, 0, true); using MemoryStream stream = new(); HeifEncoder encoder = new() { CompressionMethod = HeifCompressionMethod.Av1, - Lossless = true + Lossless = true, + Effort = 0 }; - Assert.Throws(() => image.Save(stream, encoder)); - Assert.Equal(0, stream.Length); + image.Save(stream, encoder); + byte[] file = stream.ToArray(); + Span colorPayload = GetItemPayload(file, 1); + using Av1Decoder colorDecoder = new(Configuration.Default); + using Image colorImage = colorDecoder.Decode(colorPayload); + ObuSequenceHeader colorSequenceHeader = Assert.IsType(colorDecoder.SequenceHeader); + ObuFrameHeader colorFrameHeader = Assert.IsType(colorDecoder.FrameHeader); + + Assert.Equal(Av1ColorFormat.Yuv444, colorSequenceHeader.ColorConfig.GetColorFormat()); + Assert.Equal(0, colorFrameHeader.QuantizationParameters.BaseQIndex); + Assert.True(colorFrameHeader.CodedLossless); + Assert.True(colorFrameHeader.AllLossless); + Assert.Equal(Av1TransformMode.Only4x4, colorFrameHeader.TransformMode); + + Span alphaPayload = GetItemPayload(file, 2); + using Av1Decoder alphaDecoder = new(Configuration.Default); + using Image alphaImage = alphaDecoder.Decode(alphaPayload); + ObuFrameHeader alphaFrameHeader = Assert.IsType(alphaDecoder.FrameHeader); + Assert.Equal(0, alphaFrameHeader.QuantizationParameters.BaseQIndex); + Assert.True(alphaFrameHeader.CodedLossless); + + stream.Position = 0; + using Image decoded = Image.Load(stream); + Assert.Empty(ImageComparer.Exact.CompareImages(image, decoded)); + + string outputDirectory = Path.Combine( + TestEnvironment.ActualOutputDirectoryFullPath, + "Formats", + "Heif", + "Av1"); + + Directory.CreateDirectory(outputDirectory); + File.WriteAllBytes(Path.Combine(outputDirectory, "encoder-public-lossless-color.obu"), colorPayload.ToArray()); + File.WriteAllBytes(Path.Combine(outputDirectory, "encoder-public-lossless-alpha.obu"), alphaPayload.ToArray()); } [Fact]