diff --git a/HEIF_IMPLEMENTATION_PLAN.md b/HEIF_IMPLEMENTATION_PLAN.md index 30f443cf41..d39000f5f0 100644 --- a/HEIF_IMPLEMENTATION_PLAN.md +++ b/HEIF_IMPLEMENTATION_PLAN.md @@ -16,8 +16,8 @@ This plan is the authoritative delivery checklist. A source file, unit test, bui Reference checkout evidence on 2026-08-31: -- `D:\GitHub\AOMediaCodec\aom` is clean. Its checked-out `HEAD` is `441c439b9916474cac15d2822af47a9ad70674a8`, while the refreshed `origin/main` is `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727`. -- Current encoder verification uses the exact `origin/main` tree exported to `D:\GitHub\ynse01\aom-main-a40` and built in `D:\GitHub\ynse01\aom-main-a40-build`. The resulting `aomdec` identifies itself as version 3.15.0. This records the tree audited on that date; it is not a pin and must not prevent later work from updating to the then-current `main`. +- `D:\GitHub\AOMediaCodec\aom` is clean. Its checked-out `HEAD` is `441c439b9916474cac15d2822af47a9ad70674a8`, while the refreshed `origin/main` used for the current source comparison is `d773924c767f7433d3838dbe1ba90978b10bf675`. +- The most recent independently built reference decoder uses the earlier `origin/main` tree exported to `D:\GitHub\ynse01\aom-main-a40` and built in `D:\GitHub\ynse01\aom-main-a40-build` at `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727`. The resulting `aomdec` identifies itself as version 3.15.0. This records the binary used for the last external acceptance pass; source comparison continues against the current revision above. ## Status notation @@ -825,13 +825,13 @@ Encoder verification contract: - [~] A production single-tile all-intra writer now walks raster superblocks, analyzes each immediately before entropy coding, reuses one decision workspace and one block workspace, and retains decoder-identical reconstructed references across the tile. Its closed byte and high-bit-depth operators feed the existing superblock boundary without runtime sample-type checks. A byte-exact test compares this composed path with an explicit superblock-then-tile-writer oracle, so producer and writer traversal or coefficient-area drift cannot pass unnoticed. A separate clipped 2x2-superblock regression proves global raster indexing by requiring all four coefficient segments and the bottom-right reconstruction to be populated. Multi-tile ownership and the complete frame/OBU operation remain. - [~] A non-owning encoder-frame view now separates visible conversion regions from coded regions and performs complete left, top, right, bottom, and corner extension across each bordered plane. Current libaom uses 8-sample-aligned coded dimensions, a 32-sample-aligned luma stride with chroma stride derived from it, and a 64-pixel luma border for non-resized all-intra encoding. One operation-ready frame owner now rents the aligned Y, U, and V storage contiguously, exposes non-owning `Buffer2D` plane views, and returns the rent exactly once. A 4K 4:2:0 frame occupies about 13.0 MiB at 8-bit or 26.0 MiB at 10/12-bit; source and reconstruction therefore remain distinct frame owners rather than adding a full-frame copy. The corrected tests use this real ownership path and verify the exact 54 KiB 64x64 4:2:0 rent. The frame-encoder operation now instantiates matching source and reconstruction owners with ordinary `using` lifetimes and converts packed pixels directly into the source owner before extension. - [~] Temporal delimiter, sequence header, frame header, combined-frame tile-group writing, and an internal reduced-still-picture frame operation now exist locally. The remaining required metadata, padding, multi-tile, option, and public encoder paths are not complete. -- [~] Implement superblock and partition analysis for every permitted block size and partition. Efforts zero through eight deliberately split every in-frame node to 8x8 blocks. Efforts nine and ten now perform live rate-distortion selection among `PARTITION_NONE`, `PARTITION_SPLIT`, `PARTITION_HORZ`, and `PARTITION_VERT` at each 8x8 node in current-libaom search order. Split trials visit four 4x4 leaves in raster order, rectangular trials visit two 8x4 or 4x8 leaves, and every non-final trial leaf publishes reconstructed mode, transform, and coefficient contexts for the next leaf without mutating entropy distributions. The reusable block owner grew by four integer elements for exact parent-edge restoration; no partition trial allocates or copies probability state. Production picture state retains the compact 8x8 mode allocation below effort nine and explicitly selects 4x4 allocation granularity when sub-8x8 partitions are enabled. Partition selection above 8x8, asymmetric and four-way partitions, and effort-dependent pruning remain. +- [~] Implement superblock and partition analysis for every permitted block size and partition. Efforts zero through eight deliberately split every in-frame node to 8x8 blocks. Efforts nine and ten now perform recursive live rate-distortion selection at complete 8x8 and 16x16 nodes. Candidate order matches current libaom: `PARTITION_NONE`, `PARTITION_SPLIT`, `PARTITION_HORZ`, `PARTITION_VERT`, the four asymmetric partitions, then `PARTITION_HORZ_4` and `PARTITION_VERT_4`. Invalid chroma geometries are excluded before evaluation. Each candidate saves and restores the exact partition, coefficient, transform, and palette neighbor edges in one aligned block-workspace owner; trials neither allocate nor copy probability state. Recursive split trials publish each selected child's decoded mode, transform, coefficient, and palette contexts before evaluating its next sibling. Coefficient contexts are published per retained transform rather than broadcasting the first transform over an entire partition leaf. Production picture state retains the compact 8x8 mode allocation below effort nine and explicitly selects 4x4 allocation granularity when sub-8x8 partitions are enabled. Partition selection above 16x16 and effort-dependent pruning remain. - [~] Implement intra mode search, palette, filter intra, chroma-from-luma, and intra-block copy decisions. Live luma search now covers all 13 zero-angle base modes and all six nonzero adjustments for each of the eight directional modes. Joint spatial chroma search covers the same 61 candidates, combines both chroma planes in one rate-distortion decision, and preserves the winning shared angle adjustment. Chroma-from-luma now searches the complete signed alpha alphabet from reconstructed luma and retains its joint U/V syntax. Filter-intra now searches all five predictors after ordinary luma modes. Palette entropy, retained state, production syntax, exhaustive luma and paired chroma palette selection, adaptive screen-content activation, and joint intra-block-copy mode selection are complete. - [ ] Implement inter mode search for bounded sequences, including reference selection and the decoder-supported inter tools. - [~] Current-libaom `av1_quantize_fp_no_qmatrix` arithmetic is implemented as a closed generic forward-quantizer family with Vector512, Vector256, Vector128, and scalar paths, raster-order output, coded 64-point coefficient limits, and scan-order EOB selection. Transform search, coefficient optimization, and lossless behavior remain. -- [~] Implement real rate-distortion selection and make quality and effort change work, size, and output quality. The complete luma and joint chroma candidate sets, including chroma-from-luma, filter-intra, palette, and intra-block copy, now perform live rate-distortion selection. Public quality mapping and effort tiers through exhaustive uniform luma mode/transform search are implemented. Efforts nine and ten add exact 8x8 partition rate-distortion selection; partition search above 8x8, effort-dependent pruning, and the remaining sequence searches remain. -- [~] Frame effort now progressively expands the available current search: zero is DC-only, one adds every zero-angle spatial mode, two adds every legal directional adjustment, three refines the preliminary luma winner's transform type, four adds filter-intra and chroma-from-luma, and five adds adaptive palette and intra-block-copy analysis. Lower tiers do not signal unavailable sequence or frame tools, and tiers below five skip the whole-frame screen-content scan. Effort six enables `TX_MODE_SELECT` and compares the winning ordinary spatial or filter-intra luma mode as one 8x8 transform against four raster-ordered 4x4 transforms; each luma palette candidate owns that size comparison from effort six onward. Effort seven searches every legal 8x8 transform type inside every ordinary spatial candidate rather than refining only the preliminary winner. Effort eight also performs the 8x8-versus-four-4x4 comparison inside every ordinary spatial and filter-intra candidate, matching current libaom's per-candidate uniform-transform ownership. Efforts nine and ten additionally search every legal 8x8 partition in current-libaom order. Prediction and residual construction run once per mode and are reused across its legal transform types. Every 4x4 transform searches all legal types with live coefficient contexts and reconstructed intra references. The search reuses the aligned block workspace, preserves only global improvements, and performs no per-block, per-partition, or per-transform rent. Non-skipped intra-block copy writes and costs the current-libaom unsplit variable-transform root; skipped intra-block copy emits no transform-partition symbol. Partition search above 8x8 and effort-dependent model/transform pruning remain. Decoder-visible production cases now execute effort zero through eight and ten, inspect the emitted restrictions and frame state, and decode the produced streams. The complete non-HEVC HEIF/AV1 namespace passes 9,301 of 9,301 through one foreground net11 VSTest run. Current-main `aomdec` at `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727` accepts the generated effort-eight and effort-ten payloads as well as the existing palette and intra-block-copy payloads. The affected partition and mode-decision surface passes 231 of 231 through one foreground net11 Release VSTest run, including a production effort-nine stream that selects sub-8x8 rectangular blocks and decodes successfully. The net11 Release build and Roslynk compiler pass report zero errors. -- [~] Encoder rate accounting converts the entropy writer's live inverse cumulative distributions into current-libaom fixed-point symbol costs without allocating or duplicating probability state. Read-only luma-mode, directional-delta, filter-intra, chroma-mode, block-skip, transform-size, transform-block-skip, and complete transform-coefficient queries share the exact distributions mutated by the subsequent entropy write. Complete coefficient costing follows current libaom's optimized shape: it returns immediately for an empty transform, uses the EOB-specific base-range context, fuses magnitude, sign, base-range, and Golomb accounting into one reverse traversal, and combines repeated full base-range chunks instead of replaying each emitted symbol. Tile-lifetime level and context scratch is reused, the one-coefficient path neither clears nor initializes the forward-neighbor level map, and steady-state queries allocate nothing. Transform-size writing and costing share one subdivision-depth calculation, while shared closed symbol operations keep the writer and cost mappings for transform skip, transform type, and EOB syntax identical without forcing the estimator through the writer's slower two-pass coefficient traversal. The current-libaom fixed-point RD combiner preserves 64-bit distortion and rounds the weighted 1/512-bit rate at the required boundary. Its key-frame multiplier follows libaom's squared DC-quantizer formula and exact 10/12-bit normalization. Live final-block selection evaluates all 61 legal 8x8 luma candidates: the 13 zero-angle base modes in current-libaom order, followed by six nonzero adjustments for each directional mode. Joint chroma selection evaluates the equivalent 61 spatial candidates, combines U and V distortion plus coefficient rate, and charges one live chroma-mode and shared-angle symbol over the actual subsampled 4x4, 4x8, or 8x8 geometry. Chroma-from-luma subsamples the reconstructed luma block once into fixed-stride Q3 stack scratch, subtracts the rounded mean, evaluates all 33 signed alpha values independently for each plane with complete transform RD, and combines the cached plane results across all 1,088 valid joint pairs with one live sign cost and the conditional U/V magnitude costs. This is the allocation-free equivalent of current libaom's exhaustive 33-value path: it requires 66 evaluation transforms rather than transforming every joint pair, preserves DC-before-CfL-before-spatial tie order, and fixes the implicit chroma transform to DCT-DCT. Filter-intra follows ordinary luma candidates, searches all five predictors in syntax order, and evaluates every legal transform while reusing one prepared prediction and source residual per filter mode. Every candidate includes its live mode, angle, filter mode, alpha, and coefficient rate plus normalized pixel-domain distortion. Each prepared reference edge retains the common-corner prefix and twice the transform dimension required by directional prediction. A shared encoder/decoder availability calculation selects reconstructed top-right and bottom-left extensions according to tile, frame, superblock, and block reconstruction order; unavailable extensions repeat the nearest coded endpoint. Missing top or left edges retain current libaom's perpendicular-sample and bit-depth-midpoint rules. Directional prediction applies the AV1 three-degree adjustment step and reuses transform workspace for zone-three transposition before the transform overwrites it, keeping candidate evaluation allocation-free. The winning luma and chroma signed adjustments are retained in the packed final-block state consumed by the tile writer. The tile writer invokes these reusable workspace-backed selectors after mapping current neighbors and immediately before writing each block, so later decisions see reconstructed samples, coefficient contexts, and CDF updates from every preceding block. Block skip is read only after the callback has combined every coded plane. Luma and chroma candidate scratch is partitioned from the encoder's single aligned reusable block workspace; transform-size search uses that owner for four retained 4x4 transform states, local coefficient contexts, and the compact trial reconstruction needed to preserve the best result. No candidate path rents a buffer per block or per transform. Only a newly winning candidate is copied into retained frame storage. Production fixtures force every luma base predictor, both extreme adjustments in all three directional zones, available top-right and bottom-left extensions, high-bit-depth adjustment propagation, exact signed luma and chroma angle-rate terms, joint U/V decisions, packed chroma state, and 4:2:0, 4:2:2, and 4:4:4 transform geometry. The CfL fixtures derive target chroma from a pilot production encode's actual reconstructed luma through an independent scalar Q3 oracle and prove exact positive/negative alpha syntax plus zero-residual DCT-DCT reconstruction for all three subsampling geometries at 8, 10, and 12 bits. The stable fixed-DC traversal comparison uses neutral samples for which both the baseline and live search are contractually DC and skipped, instead of relying on textured content to happen to select the baseline mode. Luma palette selection now evaluates dominant-color and one-dimensional K-means candidates for every legal size, snaps near-cache colors with the reference threshold and tie order, removes duplicate snapped colors, extends boundary maps from active samples, and performs complete transform rate-distortion search. Ordinary DC and filter-intra candidates pay the palette-disabled symbol whenever screen-content syntax is enabled. The exact net11 Release rebuild reports 1,992 test-project warnings and zero errors, all 58 intra-superblock cases pass, all 8,935 AVIF cases pass, and all 230 HEIF cases pass. Remaining mode decision work includes transform-size coverage for filter-intra and palette, broader joint mode/transform refinement, partition search above 8x8, extended partition shapes, and effort-dependent pruning. Non-empty intra blocks deliberately remain non-skipped, matching current libaom; later inter mode selection owns its distinct skip-transform RD decision. +- [~] Implement real rate-distortion selection and make quality and effort change work, size, and output quality. The complete luma and joint chroma candidate sets, including chroma-from-luma, filter-intra, palette, and intra-block copy, now perform live rate-distortion selection. Public quality mapping and effort tiers through exhaustive uniform luma mode/transform search are implemented. Efforts nine and ten add exact recursive 8x8 and 16x16 partition rate-distortion selection; partition search above 16x16, effort-dependent pruning, and the remaining sequence searches remain. +- [~] Frame effort now progressively expands the available current search: zero is DC-only, one adds every zero-angle spatial mode, two adds every legal directional adjustment, three refines the preliminary luma winner's transform type, four adds filter-intra and chroma-from-luma, and five adds adaptive palette and intra-block-copy analysis. Lower tiers do not signal unavailable sequence or frame tools, and tiers below five skip the whole-frame screen-content scan. Effort six enables `TX_MODE_SELECT` and compares the winning ordinary spatial or filter-intra luma mode as one 8x8 transform against four raster-ordered 4x4 transforms; each luma palette candidate owns that size comparison from effort six onward. Effort seven searches every legal 8x8 transform type inside every ordinary spatial candidate rather than refining only the preliminary winner. Effort eight also performs the 8x8-versus-four-4x4 comparison inside every ordinary spatial and filter-intra candidate, matching current libaom's per-candidate uniform-transform ownership. Efforts nine and ten additionally search every legal partition at complete 8x8 and 16x16 nodes in current-libaom order. Prediction and residual construction run once per mode and are reused across its legal transform types. Every 4x4 transform searches all legal types with live coefficient contexts and reconstructed intra references. The search reuses the aligned block workspace, preserves only global improvements, and performs no per-block, per-partition, or per-transform rent. Non-skipped intra-block copy writes and costs the current-libaom unsplit variable-transform root; skipped intra-block copy emits no transform-partition symbol. Partition search above 16x16 and effort-dependent model/transform pruning remain. Decoder-visible production cases inspect the emitted restrictions and frame state and decode the produced streams, including real effort-nine streams selecting sub-8x8 and 8x16 rectangular blocks. The complete non-HEVC HEIF/AV1 namespace passes 9,071 of 9,071 through one foreground net11 Release VSTest run. The last independently built `aomdec`, from the then-current `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727` snapshot, accepts the previously generated effort-eight and effort-ten payloads as well as the existing palette and intra-block-copy payloads. The affected partition and mode-decision surface passes 232 of 232 through one foreground net11 Release VSTest run. The net11 Release build and Roslynk compiler and analyzer passes report zero errors. +- [~] Encoder rate accounting converts the entropy writer's live inverse cumulative distributions into current-libaom fixed-point symbol costs without allocating or duplicating probability state. Read-only luma-mode, directional-delta, filter-intra, chroma-mode, block-skip, transform-size, transform-block-skip, and complete transform-coefficient queries share the exact distributions mutated by the subsequent entropy write. Complete coefficient costing follows current libaom's optimized shape: it returns immediately for an empty transform, uses the EOB-specific base-range context, fuses magnitude, sign, base-range, and Golomb accounting into one reverse traversal, and combines repeated full base-range chunks instead of replaying each emitted symbol. Tile-lifetime level and context scratch is reused, the one-coefficient path neither clears nor initializes the forward-neighbor level map, and steady-state queries allocate nothing. Transform-size writing and costing share one subdivision-depth calculation, while shared closed symbol operations keep the writer and cost mappings for transform skip, transform type, and EOB syntax identical without forcing the estimator through the writer's slower two-pass coefficient traversal. The current-libaom fixed-point RD combiner preserves 64-bit distortion and rounds the weighted 1/512-bit rate at the required boundary. Its key-frame multiplier follows libaom's squared DC-quantizer formula and exact 10/12-bit normalization. Live final-block selection evaluates all 61 legal 8x8 luma candidates: the 13 zero-angle base modes in current-libaom order, followed by six nonzero adjustments for each directional mode. Joint chroma selection evaluates the equivalent 61 spatial candidates, combines U and V distortion plus coefficient rate, and charges one live chroma-mode and shared-angle symbol over the actual subsampled 4x4, 4x8, or 8x8 geometry. Chroma-from-luma subsamples the reconstructed luma block once into fixed-stride Q3 stack scratch, subtracts the rounded mean, evaluates all 33 signed alpha values independently for each plane with complete transform RD, and combines the cached plane results across all 1,088 valid joint pairs with one live sign cost and the conditional U/V magnitude costs. This is the allocation-free equivalent of current libaom's exhaustive 33-value path: it requires 66 evaluation transforms rather than transforming every joint pair, preserves DC-before-CfL-before-spatial tie order, and fixes the implicit chroma transform to DCT-DCT. Filter-intra follows ordinary luma candidates, searches all five predictors in syntax order, and evaluates every legal transform while reusing one prepared prediction and source residual per filter mode. Every candidate includes its live mode, angle, filter mode, alpha, and coefficient rate plus normalized pixel-domain distortion. Each prepared reference edge retains the common-corner prefix and twice the transform dimension required by directional prediction. A shared encoder/decoder availability calculation selects reconstructed top-right and bottom-left extensions according to tile, frame, superblock, and block reconstruction order; unavailable extensions repeat the nearest coded endpoint. Missing top or left edges retain current libaom's perpendicular-sample and bit-depth-midpoint rules. Directional prediction applies the AV1 three-degree adjustment step and reuses transform workspace for zone-three transposition before the transform overwrites it, keeping candidate evaluation allocation-free. The winning luma and chroma signed adjustments are retained in the packed final-block state consumed by the tile writer. The tile writer invokes these reusable workspace-backed selectors after mapping current neighbors and immediately before writing each block, so later decisions see reconstructed samples, coefficient contexts, and CDF updates from every preceding block. Block skip is read only after the callback has combined every coded plane. Luma and chroma candidate scratch is partitioned from the encoder's single aligned reusable block workspace; transform-size search uses that owner for four retained 4x4 transform states, local coefficient contexts, and the compact trial reconstruction needed to preserve the best result. No candidate path rents a buffer per block or per transform. Only a newly winning candidate is copied into retained frame storage. Production fixtures force every luma base predictor, both extreme adjustments in all three directional zones, available top-right and bottom-left extensions, high-bit-depth adjustment propagation, exact signed luma and chroma angle-rate terms, joint U/V decisions, packed chroma state, and 4:2:0, 4:2:2, and 4:4:4 transform geometry. The CfL fixtures derive target chroma from a pilot production encode's actual reconstructed luma through an independent scalar Q3 oracle and prove exact positive/negative alpha syntax plus zero-residual DCT-DCT reconstruction for all three subsampling geometries at 8, 10, and 12 bits. The stable fixed-DC traversal comparison uses neutral samples for which both the baseline and live search are contractually DC and skipped, instead of relying on textured content to happen to select the baseline mode. Luma palette selection now evaluates dominant-color and one-dimensional K-means candidates for every legal size, snaps near-cache colors with the reference threshold and tie order, removes duplicate snapped colors, extends boundary maps from active samples, and performs complete transform rate-distortion search. Ordinary DC and filter-intra candidates pay the palette-disabled symbol whenever screen-content syntax is enabled. The exact net11 Release rebuild reports 1,992 test-project warnings and zero errors, all 58 intra-superblock cases pass, all 8,935 AVIF cases pass, and all 230 HEIF cases pass. Remaining mode decision work includes transform-size coverage for filter-intra and palette, broader joint mode/transform refinement, partition search above 16x16 and effort-dependent pruning. Non-empty intra blocks deliberately remain non-skipped, matching current libaom; later inter mode selection owns its distinct skip-transform RD decision. - [~] The tile writer now publishes one packed coefficient context per covered 4x4 edge unit and derives luma/chroma skip plus DC-sign contexts from the complete transform edges using current-libaom units. Partition, transform, and coefficient neighbor state retains only the above and left context regions used by current libaom; the unused third top-left region, its granularity state, and its unused sentinel are removed. One picture owner now packs segmentation plus every tile's partition, luma, chroma, and transform edges into one clean byte allocation with typed non-owning views; together with the separately typed packed mode-information owner, the complete picture state uses two allocator rents rather than seven. Exact aligned lengths, clean initialization, and balanced exactly-once returns are covered in Release. Multi-tile payload ownership and verified CDF update behavior remain. - [~] Encoder mode information now uses a frame-owned integer alias grid over a packed 8-byte value allocation, matching current libaom's `mi_grid_base` and `mi_alloc` relationship without a managed object or reference per 4x4 entry. The visible dimensions are aligned to eight luma samples, the grid stride and allocated row count are aligned to 32 mode-information units, and optional 8x8 allocation granularity reduces the value store in both dimensions exactly as current libaom does. One clean ImageSharp byte owner contains both independently typed regions, reducing libaom's two allocation lifetimes to one without a copy. At 4K, the 4x4 layout occupies about 6.0 MiB in total; the 8x8 layout occupies about 3.0 MiB. Exact geometry, clean allocation, typed lengths, aligned mapping, untouched row padding, and exactly-once return pass 4 of 4 direct net11 VSTest cases in Release. Every coded 4x4 cell covered by square, rectangular, or clipped edge blocks maps to its owning allocation entry before context-dependent symbols are written. Packed syntax, relative neighbor lookup, full block mapping, writer traversal, entropy, and OBU coverage pass 1,947 of 1,947 direct net11 VSTest cases in Release; complete mode decision still remains. - [~] The final-block decision workspace uses one reusable 8.3 KiB ImageSharp allocator owner. It contains 1,024 explicitly packed 8-byte final-block entries and the 341 preorder partition bytes required by a complete 128x128-through-8x8 quadtree, replacing separate managed arrays. Palette colors now have their own current-block value and are copied only to the picture edges that later blocks can reference, so enabling palette mode does not add 50 bytes to every final-block entry. Construction and the explicit per-superblock reset initialize every syntax field, including the nonzero sentinel that disables filter-intra prediction; pooled quantizer, prediction, partition, and current-palette bytes cannot leak into the next decision pass. Complete mode decision still remains. @@ -868,7 +868,7 @@ Encoder verification contract: - [x] Operation-wide allocation tracking now exercises a real 64x64 12-bit 4:4:4 frame through packed-pixel conversion, both native frame owners, picture and coefficient state, reusable block workspaces, entropy coding, OBU framing, and a non-seekable destination. It proves exactly one 60 KiB tile-output reservation from current libaom's all-intra 2.5x rule and balanced exactly-once returns for every tracked allocation before the operation completes. The focused ownership case passes 1 of 1 and the complete HEIF/AV1 namespace passes 8,863 of 8,863 direct net11 VSTest cases with zero failures or skips. - [x] HEIF box offsets are now counted from the start of the encoded file instead of reading `Stream.Position`. This preserves ISO BMFF file-relative `iloc` offsets when the destination begins at a nonzero position and permits non-seekable output. Decoder item extents and image-sequence chunk offsets now resolve from that same file origin rather than the backing stream origin. Real legacy-JPEG HEIF round trips cover non-seekable output and a prefixed destination, while current-position AV1 decode covers both a still item and a five-frame sequence. All 96 encoder/decoder cases and all 38 sequence-parser cases pass direct net11 Release VSTest; the Release build remains at the established 1,005-warning baseline with zero errors. -- [~] Current-libaom source comparison now drives uniform luma transform ownership at each effort boundary. Effort six retains the cheaper winner-only size decision for ordinary spatial and filter-intra modes, while every palette candidate already owns its size decision. Effort seven evaluates every legal 8x8 transform type for every ordinary spatial candidate. Effort eight and above make transform size part of every ordinary spatial and filter-intra candidate's rate-distortion result, so an 8x8-only preliminary comparison cannot discard the mode that wins with four 4x4 transforms. Prediction and subtraction are prepared once per mode and reused across transform types, matching the reference separation between prediction and transform search. Each 4x4 transform searches every legal type with coefficient contexts derived from retained transform edges and preceding trial blocks, while reconstructed top-right and bottom-left references follow production coding order. Palette prediction uses non-owning subregions of the retained color map, and filter-intra rebuilds each recursive prediction from reconstructed edges. The existing aligned block-workspace owner retains prediction, residual, coefficients, contexts, compact reconstruction, and four final states; no allocator rent, managed array, best-candidate re-transform, or full-block intermediate copy was added. Dense decision points now document scratch lifetime, enumeration tie order, global-winner publication, raster reconstruction dependencies, and the deliberate lower-effort shortcut. The packed encoder transform edges initialize to 64, matching libaom and the ImageSharp decoder before a coded neighbor publishes its size, and variable transform syntax remains gated to blocks larger than 4x4. The focused Release verification passes 13 of 13 cases across efforts zero through eight and ten, palette split selection, and transform-size selection. The complete non-HEVC HEIF/AV1 namespace passes 9,301 of 9,301 cases with zero failures or skips. Current-main `aomdec` at `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727` accepts the generated effort-eight and effort-ten streams. Partition search and effort-dependent pruning remain. +- [~] Current-libaom source comparison now drives uniform luma transform ownership at each effort boundary. Effort six retains the cheaper winner-only size decision for ordinary spatial and filter-intra modes, while every palette candidate already owns its size decision. Effort seven evaluates every legal 8x8 transform type for every ordinary spatial candidate. Effort eight and above make transform size part of every ordinary spatial and filter-intra candidate's rate-distortion result, so an 8x8-only preliminary comparison cannot discard the mode that wins with four 4x4 transforms. Prediction and subtraction are prepared once per mode and reused across transform types, matching the reference separation between prediction and transform search. Each 4x4 transform searches every legal type with coefficient contexts derived from retained transform edges and preceding trial blocks, while reconstructed top-right and bottom-left references follow production coding order. Palette prediction uses non-owning subregions of the retained color map, and filter-intra rebuilds each recursive prediction from reconstructed edges. The existing aligned block-workspace owner retains prediction, residual, coefficients, contexts, compact reconstruction, and four final states; no allocator rent, managed array, best-candidate re-transform, or full-block intermediate copy was added. Dense decision points now document scratch lifetime, enumeration tie order, global-winner publication, raster reconstruction dependencies, and the deliberate lower-effort shortcut. The packed encoder transform edges initialize to 64, matching libaom and the ImageSharp decoder before a coded neighbor publishes its size, and variable transform syntax remains gated to blocks larger than 4x4. The focused Release verification passes 13 of 13 cases across efforts zero through eight and ten, palette split selection, and transform-size selection. The complete non-HEVC HEIF/AV1 namespace passed 9,301 of 9,301 cases at that checkpoint. The `aomdec` built from the then-current `a40ed1ea9e4ecc3df58a5bccb76623f2c94ae727` snapshot accepts the generated effort-eight and effort-ten streams. Partition search and effort-dependent pruning remain. - [x] Intra-block-copy transform search now prepares motion compensation and subtraction once per plane, alternates the existing candidate and selected work buffers whenever a transform improves, and performs at most one final normalization copy into the caller-owned selected span. This matches current libaom's pointer-swap ownership without adding an allocation or a third reconstruction buffer. Inline documentation now records the scratch lifetime, strict transform tie order, skip-rate replacement, unsplit transform-root syntax, joint-plane winner retention, and final publication boundary. The focused Release encoder and intra-block-copy set passes 17 of 17 cases, the complete non-HEVC HEIF/AV1 namespace passes 9,301 of 9,301 cases with zero failures or skips, and current-main `aomdec` accepts the regenerated effort-five and effort-six intra-block-copy streams. diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderBlockWorkspace.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderBlockWorkspace.cs index 20f4d37d2e..9386e92e7d 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderBlockWorkspace.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderBlockWorkspace.cs @@ -3,6 +3,7 @@ using System.Buffers; using System.Runtime.InteropServices; +using SixLabors.ImageSharp.Formats.Heif.Av1.Tiling; using SixLabors.ImageSharp.Formats.Heif.Av1.Transform; using SixLabors.ImageSharp.Memory; @@ -31,7 +32,7 @@ internal sealed class Av1EncoderBlockWorkspace : IDisposable MaximumCoefficientCount + MaximumCoefficientCount + Av1TransformWorkspace.MaximumLength + - IntraBlockCopyStorageLength + + SharedModeDecisionStorageLength + PartitionContextStorageLength; private const int ResidualStorageLength = MaximumResidualCount / 2; @@ -65,10 +66,31 @@ internal sealed class Av1EncoderBlockWorkspace : IDisposable IntraBlockCopyResidualStorageLength + IntraBlockCopyCoefficientStorageLength; + private const int ModeDecisionStorageLength = Av1EncoderModeDecisionWorkspace.StorageLength; + private const int SharedModeDecisionStorageLength = ModeDecisionStorageLength > IntraBlockCopyStorageLength + ? ModeDecisionStorageLength + : IntraBlockCopyStorageLength; + private const int PartitionContextStorageOffset = - IntraBlockCopySampleStorageOffset + IntraBlockCopyStorageLength; + IntraBlockCopySampleStorageOffset + SharedModeDecisionStorageLength; + + private const int MaximumPartitionEdgeUnitCount = + 2 * (1 << (Av1Constants.MaxSuperBlockSizeLog2 - Av1Constants.ModeInfoSizeLog2)); + + private const int PartitionContextBytesPerEdgeUnit = + Av1PartitionContext.StorageSize + (4 * sizeof(byte)) + Av1EncoderPaletteInfo.StorageSize; + + private const int PartitionContextSlotByteLength = + MaximumPartitionEdgeUnitCount * PartitionContextBytesPerEdgeUnit; + + private const int PartitionContextSlotLength = + PartitionContextSlotByteLength / sizeof(int); - private const int PartitionContextStorageLength = 4; + private const int PartitionTrialLevelCount = + Av1Constants.MaxSuperBlockSizeLog2 - 3 + 1; + + private const int PartitionContextStorageLength = + PartitionContextSlotLength * PartitionTrialLevelCount; /// /// Owns the complete reusable block workspace in 32-bit elements so every transform region is naturally aligned. @@ -107,13 +129,20 @@ internal sealed class Av1EncoderBlockWorkspace : IDisposable => this.owner.Memory.Span.Slice(TransformWorkspaceOffset, Av1TransformWorkspace.MaximumLength); /// - /// Gets storage for the coefficient and transform edges restored after an 8x8 partition trial. + /// Gets the disjoint edge snapshot used to restore one square partition-search level. /// - public Span PartitionContexts - => MemoryMarshal.AsBytes( - this.owner.Memory.Span.Slice( - PartitionContextStorageOffset, - PartitionContextStorageLength)); + /// The square partition node being evaluated. + /// The maximum-size byte view reserved for that node depth. + public Span GetPartitionContextStorage(Av1BlockSize blockSize) + { + int blockSizeLog2 = Av1Math.Log2(blockSize.GetWidth()); + int slotIndex = Av1Constants.MaxSuperBlockSizeLog2 - blockSizeLog2; + Span storage = this.owner.Memory.Span.Slice( + PartitionContextStorageOffset + (slotIndex * PartitionContextSlotLength), + PartitionContextSlotLength); + + return MemoryMarshal.AsBytes(storage); + } /// /// Gets the reusable storage used while comparing spatial, chroma-from-luma, filter-intra, and palette candidates. @@ -127,7 +156,7 @@ internal sealed class Av1EncoderBlockWorkspace : IDisposable // Both phases can therefore reuse this aligned region without extending the owner or preserving stale scratch. Span storage = this.owner.Memory.Span.Slice( IntraBlockCopySampleStorageOffset, - IntraBlockCopyStorageLength); + SharedModeDecisionStorageLength); return new(storage[..Av1EncoderModeDecisionWorkspace.StorageLength]); } diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs index f93708151d..fbbd50fdad 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1EncoderModeDecisionWorkspace.cs @@ -16,9 +16,14 @@ internal readonly ref struct Av1EncoderModeDecisionWorkspace where TSample : unmanaged { /// - /// The maximum number of samples in the encoder's fixed 8x8 transform block. + /// The largest coding-block dimension evaluated directly by the current partition search. /// - public const int MaximumSampleCount = 8 * 8; + public const int MaximumBlockDimension = 16; + + /// + /// The maximum number of samples in one directly evaluated coding block. + /// + public const int MaximumSampleCount = MaximumBlockDimension * MaximumBlockDimension; /// /// The number of 4x4 transform blocks covering one 8x8 coding block. @@ -30,7 +35,7 @@ internal readonly ref struct Av1EncoderModeDecisionWorkspace /// public const int StorageLength = TransientStorageOffset + Av1EncoderPaletteWorkspace.StorageLength; - private const int ReferenceBufferLength = 17; + private const int ReferenceBufferLength = (2 * MaximumBlockDimension) + 1; private const int ReferenceBufferCount = 4; private const int ReferenceStorageLength = ReferenceBufferCount * ReferenceBufferLength * sizeof(ushort) / sizeof(int); private const int CandidateSampleStorageOffset = ReferenceStorageLength; @@ -40,9 +45,13 @@ internal readonly ref struct Av1EncoderModeDecisionWorkspace private const int CandidateTransformBlockStorageOffset = CandidateCoefficientStorageOffset + CandidateCoefficientStorageLength; private const int CandidateTransformBlockStorageLength = CandidateTransformBlockCount; private const int TransformContextStorageOffset = CandidateTransformBlockStorageOffset + CandidateTransformBlockStorageLength; - private const int TransformContextStorageLength = 1; + private const int TransformContextStorageLength = + 2 * (MaximumBlockDimension >> Av1Constants.ModeInfoSizeLog2) * sizeof(byte) / sizeof(int); + private const int TransientStorageOffset = TransformContextStorageOffset + TransformContextStorageLength; - private const int ChromaFromLumaSampleCount = Av1ChromaFromLumaContext.BufferLine * 8; + private const int ChromaFromLumaSampleCount = + Av1ChromaFromLumaContext.BufferLine * MaximumBlockDimension; + private const int ChromaFromLumaSampleStorageLength = ChromaFromLumaSampleCount * sizeof(short) / sizeof(int); private const int ChromaFromLumaBlueRateOffset = ChromaFromLumaSampleStorageLength; private const int ChromaFromLumaRedRateOffset = ChromaFromLumaBlueRateOffset + Av1ChromaFromLumaMath.AlphaCandidateCount; diff --git a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs index 157c441cda..b65f04d801 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Pipeline/Av1IntraSuperblockEncoder.ModeDecision.cs @@ -1,6 +1,7 @@ // Copyright (c) Six Labors. // Licensed under the Six Labors Split License. +using System.Runtime.InteropServices; using SixLabors.ImageSharp.Formats.Heif.Av1.Entropy; using SixLabors.ImageSharp.Formats.Heif.Av1.OpenBitstreamUnit; using SixLabors.ImageSharp.Formats.Heif.Av1.Prediction; @@ -40,6 +41,23 @@ internal static partial class Av1IntraSuperblockEncoder /// private static ReadOnlySpan AngleDeltaSearchOrder => [-3, -2, -1, 1, 2, 3]; + /// + /// Gets partition candidates in the evaluation order used by the reference encoder. + /// + private static ReadOnlySpan PartitionSearchOrder => + [ + Av1PartitionType.None, + Av1PartitionType.Split, + Av1PartitionType.Horizontal, + Av1PartitionType.Vertical, + Av1PartitionType.HorizontalA, + Av1PartitionType.HorizontalB, + Av1PartitionType.VerticalA, + Av1PartitionType.VerticalB, + Av1PartitionType.Horizontal4, + Av1PartitionType.Vertical4 + ]; + /// /// Builds the fixed 8x8 partition skeleton consumed by interleaved mode decision and tile writing. /// @@ -165,32 +183,24 @@ internal static partial class Av1IntraSuperblockEncoder Av1BlockSize blockSize, Av1PartitionType preparedPartition) { - if (blockSize != Av1BlockSize.Block8x8) + if (blockSize is not Av1BlockSize.Block8x8 and not Av1BlockSize.Block16x16) { return preparedPartition; } Point modeInfoPosition = blockOrigin >> Av1Constants.ModeInfoSizeLog2; - bool hasRows = modeInfoPosition.Y + 1 < this.picture.Parent.Common.ModeInfoRowCount; - bool hasColumns = modeInfoPosition.X + 1 < this.picture.Parent.Common.ModeInfoColumnCount; - if (!hasRows || !hasColumns) - { - // A clipped 8x8 node must split because its missing half cannot be represented by PARTITION_NONE. - for (int childIndex = 0; childIndex < 4; childIndex++) - { - Point childOrigin = blockOrigin + new Size( - (childIndex & 1) << Av1Constants.ModeInfoSizeLog2, - (childIndex >> 1) << Av1Constants.ModeInfoSizeLog2); + bool hasRows = + modeInfoPosition.Y + blockSize.Get4x4HighCount() <= this.picture.Parent.Common.ModeInfoRowCount; - Point childPosition = childOrigin >> Av1Constants.ModeInfoSizeLog2; - if (childPosition.Y < this.picture.Parent.Common.ModeInfoRowCount && - childPosition.X < this.picture.Parent.Common.ModeInfoColumnCount) - { - this.SetBlockGeometry(childOrigin, Av1BlockSize.Block4x4, Av1PartitionType.None); - } - } + bool hasColumns = + modeInfoPosition.X + blockSize.Get4x4WideCount() <= this.picture.Parent.Common.ModeInfoColumnCount; - return Av1PartitionType.Split; + if (!hasRows || !hasColumns) + { + // Coded dimensions are aligned to eight samples, so an incomplete searched node must retain + // the prepared split tree rather than evaluating a block that extends beyond source storage. + this.PreparePartitionGeometry(blockOrigin, blockSize, preparedPartition); + return preparedPartition; } if (this.effort < 9) @@ -198,68 +208,102 @@ internal static partial class Av1IntraSuperblockEncoder return preparedPartition; } - int savedLumaArea = this.codedAreaLuma; - int savedChromaArea = this.codedAreaChroma; - this.SavePartitionTrialContexts(blockOrigin, tileIndex); - long bestCost = this.EvaluatePartitionCandidate( + Av1PartitionType selectedPartition = this.SelectBestPartition( writer, macroBlock, blockOrigin, tileIndex, - blockSize, - Av1PartitionType.None); + blockSize); - Av1PartitionType selectedPartition = Av1PartitionType.None; - this.ResetPartitionTrial(blockOrigin, tileIndex, savedLumaArea, savedChromaArea); - long splitCost = this.EvaluatePartitionCandidate( - writer, - macroBlock, - blockOrigin, - tileIndex, - blockSize, - Av1PartitionType.Split); + // Trial reconstruction and mode entries need no copy-back. The selected branch is evaluated again + // in raster order, overwriting each trial-local value before a later selected leaf can consume it. + this.PreparePartitionGeometry(blockOrigin, blockSize, selectedPartition); + return selectedPartition; + } - if (splitCost < bestCost) + private Av1PartitionType SelectBestPartition( + Av1SymbolEncoder writer, + Av1MacroBlockD macroBlock, + Point blockOrigin, + ushort tileIndex, + Av1BlockSize blockSize) + { + int savedLumaArea = this.codedAreaLuma; + int savedChromaArea = this.codedAreaChroma; + this.SavePartitionTrialContexts(blockOrigin, tileIndex, blockSize); + long bestCost = long.MaxValue; + Av1PartitionType selectedPartition = Av1PartitionType.None; + ReadOnlySpan searchOrder = PartitionSearchOrder; + int candidateCount = blockSize == Av1BlockSize.Block8x8 ? 4 : searchOrder.Length; + for (int candidateIndex = 0; candidateIndex < candidateCount; candidateIndex++) { - bestCost = splitCost; - selectedPartition = Av1PartitionType.Split; + Av1PartitionType partitionType = searchOrder[candidateIndex]; + if (!this.IsPartitionCandidateAllowed(blockSize, partitionType)) + { + continue; + } + + long candidateCost = this.EvaluatePartitionCandidate( + writer, + macroBlock, + blockOrigin, + tileIndex, + blockSize, + partitionType, + publishFinalContexts: false); + + if (candidateCost < bestCost) + { + bestCost = candidateCost; + selectedPartition = partitionType; + } + + this.ResetPartitionTrial( + blockOrigin, + tileIndex, + blockSize, + savedLumaArea, + savedChromaArea); } - this.ResetPartitionTrial(blockOrigin, tileIndex, savedLumaArea, savedChromaArea); - long horizontalCost = this.EvaluatePartitionCandidate( + return selectedPartition; + } + + private long EvaluateSelectedPartitionTree( + Av1SymbolEncoder writer, + Av1MacroBlockD macroBlock, + Point blockOrigin, + ushort tileIndex, + Av1BlockSize blockSize, + bool publishContexts) + { + Av1PartitionType selectedPartition = this.SelectBestPartition( writer, macroBlock, blockOrigin, tileIndex, - blockSize, - Av1PartitionType.Horizontal); - - if (horizontalCost < bestCost) - { - bestCost = horizontalCost; - selectedPartition = Av1PartitionType.Horizontal; - } + blockSize); - this.ResetPartitionTrial(blockOrigin, tileIndex, savedLumaArea, savedChromaArea); - long verticalCost = this.EvaluatePartitionCandidate( + long cost = this.EvaluatePartitionCandidate( writer, macroBlock, blockOrigin, tileIndex, blockSize, - Av1PartitionType.Vertical); + selectedPartition, + publishContexts); - if (verticalCost < bestCost) + if (publishContexts) { - selectedPartition = Av1PartitionType.Vertical; + Av1TileWriter.UpdatePartitionContexts( + this.picture.PartitionContexts[tileIndex], + blockOrigin, + selectedPartition.GetBlockSubSize(blockSize), + blockSize, + selectedPartition); } - this.ResetPartitionTrial(blockOrigin, tileIndex, savedLumaArea, savedChromaArea); - - // Trial reconstruction and mode entries need no copy-back. The selected branch is evaluated again - // in raster order, overwriting each trial-local value before a later selected leaf can consume it. - this.PreparePartitionGeometry(blockOrigin, blockSize, selectedPartition); - return selectedPartition; + return cost; } private long EvaluatePartitionCandidate( @@ -268,7 +312,8 @@ internal static partial class Av1IntraSuperblockEncoder Point blockOrigin, ushort tileIndex, Av1BlockSize blockSize, - Av1PartitionType partitionType) + Av1PartitionType partitionType, + bool publishFinalContexts) { int rate = Av1TileWriter.GetPartitionCost( this.picture, @@ -279,31 +324,36 @@ internal static partial class Av1IntraSuperblockEncoder this.picture.PartitionContexts[tileIndex]); long cost = Av1RateDistortion.GetCost(this.rateMultiplier, rate, 0); - int leafCount = partitionType == Av1PartitionType.Split - ? 4 - : partitionType == Av1PartitionType.None - ? 1 - : 2; - - Av1BlockSize leafSize = partitionType.GetBlockSubSize(blockSize); + int leafCount = GetPartitionLeafCount(partitionType); // Child reconstruction and syntax contexts become input to the next child. Publishing only - // non-final leaves reproduces libaom's raster trial without writing entropy symbols. + // the required leaves reproduces libaom's raster dry run without writing entropy symbols. for (int leafIndex = 0; leafIndex < leafCount; leafIndex++) { - Point leafOrigin = GetPartitionLeafOrigin( + GetPartitionLeafGeometry( blockOrigin, blockSize, partitionType, - leafIndex); + leafIndex, + out Point leafOrigin, + out Av1BlockSize leafSize); - cost += this.EvaluatePartitionLeaf( - writer, - macroBlock, - leafOrigin, - tileIndex, - leafSize, - publishContexts: leafIndex < leafCount - 1); + bool publishContexts = leafIndex < leafCount - 1 || publishFinalContexts; + cost += partitionType == Av1PartitionType.Split && blockSize > Av1BlockSize.Block8x8 + ? this.EvaluateSelectedPartitionTree( + writer, + macroBlock, + leafOrigin, + tileIndex, + leafSize, + publishContexts) + : this.EvaluatePartitionLeaf( + writer, + macroBlock, + leafOrigin, + tileIndex, + leafSize, + publishContexts); } return cost; @@ -312,12 +362,13 @@ internal static partial class Av1IntraSuperblockEncoder private void ResetPartitionTrial( Point blockOrigin, ushort tileIndex, + Av1BlockSize blockSize, int savedLumaArea, int savedChromaArea) { this.codedAreaLuma = savedLumaArea; this.codedAreaChroma = savedChromaArea; - this.RestorePartitionTrialContexts(blockOrigin, tileIndex); + this.RestorePartitionTrialContexts(blockOrigin, tileIndex, blockSize); } private void PreparePartitionGeometry( @@ -325,42 +376,149 @@ internal static partial class Av1IntraSuperblockEncoder Av1BlockSize blockSize, Av1PartitionType partitionType) { - int leafCount = partitionType == Av1PartitionType.Split - ? 4 - : partitionType == Av1PartitionType.None - ? 1 - : 2; - - Av1BlockSize leafSize = partitionType.GetBlockSubSize(blockSize); + int leafCount = GetPartitionLeafCount(partitionType); for (int leafIndex = 0; leafIndex < leafCount; leafIndex++) { - Point leafOrigin = GetPartitionLeafOrigin( + GetPartitionLeafGeometry( blockOrigin, blockSize, partitionType, - leafIndex); + leafIndex, + out Point leafOrigin, + out Av1BlockSize leafSize); - this.SetBlockGeometry(leafOrigin, leafSize, Av1PartitionType.None); + if (this.IsBlockOriginInsideFrame(leafOrigin)) + { + this.SetBlockGeometry(leafOrigin, leafSize, Av1PartitionType.None); + } } } - private static Point GetPartitionLeafOrigin( + private bool IsPartitionCandidateAllowed( + Av1BlockSize blockSize, + Av1PartitionType partitionType) + { + if (partitionType.GetBlockSubSize(blockSize) == Av1BlockSize.Invalid) + { + return false; + } + + if (this.source.IsMonochrome) + { + return true; + } + + ObuColorConfig colorConfig = this.picture.Sequence.SequenceHeader.ColorConfig; + int leafCount = GetPartitionLeafCount(partitionType); + for (int leafIndex = 0; leafIndex < leafCount; leafIndex++) + { + GetPartitionLeafGeometry( + Point.Empty, + blockSize, + partitionType, + leafIndex, + out _, + out Av1BlockSize leafSize); + + if (leafSize.GetSubsampled(colorConfig.SubSamplingX, colorConfig.SubSamplingY) == + Av1BlockSize.Invalid) + { + return false; + } + } + + return true; + } + + private bool IsBlockOriginInsideFrame(Point blockOrigin) + { + Point modeInfoPosition = blockOrigin >> Av1Constants.ModeInfoSizeLog2; + return modeInfoPosition.Y < this.picture.Parent.Common.ModeInfoRowCount && + modeInfoPosition.X < this.picture.Parent.Common.ModeInfoColumnCount; + } + + private static int GetPartitionLeafCount(Av1PartitionType partitionType) + => partitionType switch + { + Av1PartitionType.None => 1, + Av1PartitionType.Horizontal or Av1PartitionType.Vertical => 2, + Av1PartitionType.HorizontalA or + Av1PartitionType.HorizontalB or + Av1PartitionType.VerticalA or + Av1PartitionType.VerticalB => 3, + _ => 4 + }; + + private static void GetPartitionLeafGeometry( Point blockOrigin, Av1BlockSize blockSize, Av1PartitionType partitionType, - int leafIndex) + int leafIndex, + out Point leafOrigin, + out Av1BlockSize leafSize) { int halfWidth = blockSize.GetWidth() >> 1; int halfHeight = blockSize.GetHeight() >> 1; - return partitionType switch + Av1BlockSize rectangularSize = partitionType.GetBlockSubSize(blockSize); + Av1BlockSize splitSize = Av1PartitionType.Split.GetBlockSubSize(blockSize); + switch (partitionType) { - Av1PartitionType.Horizontal => blockOrigin + new Size(0, leafIndex * halfHeight), - Av1PartitionType.Vertical => blockOrigin + new Size(leafIndex * halfWidth, 0), - Av1PartitionType.Split => blockOrigin + new Size( - (leafIndex & 1) * halfWidth, - (leafIndex >> 1) * halfHeight), - _ => blockOrigin - }; + case Av1PartitionType.Horizontal: + leafOrigin = blockOrigin + new Size(0, leafIndex * halfHeight); + leafSize = rectangularSize; + return; + case Av1PartitionType.Vertical: + leafOrigin = blockOrigin + new Size(leafIndex * halfWidth, 0); + leafSize = rectangularSize; + return; + case Av1PartitionType.Split: + leafOrigin = blockOrigin + new Size( + (leafIndex & 1) * halfWidth, + (leafIndex >> 1) * halfHeight); + + leafSize = splitSize; + return; + case Av1PartitionType.HorizontalA: + leafOrigin = leafIndex < 2 + ? blockOrigin + new Size(leafIndex * halfWidth, 0) + : blockOrigin + new Size(0, halfHeight); + + leafSize = leafIndex < 2 ? splitSize : rectangularSize; + return; + case Av1PartitionType.HorizontalB: + leafOrigin = leafIndex == 0 + ? blockOrigin + : blockOrigin + new Size((leafIndex - 1) * halfWidth, halfHeight); + + leafSize = leafIndex == 0 ? rectangularSize : splitSize; + return; + case Av1PartitionType.VerticalA: + leafOrigin = leafIndex < 2 + ? blockOrigin + new Size(0, leafIndex * halfHeight) + : blockOrigin + new Size(halfWidth, 0); + + leafSize = leafIndex < 2 ? splitSize : rectangularSize; + return; + case Av1PartitionType.VerticalB: + leafOrigin = leafIndex == 0 + ? blockOrigin + : blockOrigin + new Size(halfWidth, (leafIndex - 1) * halfHeight); + + leafSize = leafIndex == 0 ? rectangularSize : splitSize; + return; + case Av1PartitionType.Horizontal4: + leafOrigin = blockOrigin + new Size(0, leafIndex * (blockSize.GetHeight() >> 2)); + leafSize = rectangularSize; + return; + case Av1PartitionType.Vertical4: + leafOrigin = blockOrigin + new Size(leafIndex * (blockSize.GetWidth() >> 2), 0); + leafSize = rectangularSize; + return; + default: + leafOrigin = blockOrigin; + leafSize = blockSize; + return; + } } /// @@ -658,7 +816,8 @@ internal static partial class Av1IntraSuperblockEncoder lumaArea, chromaArea, modeInfo, - block); + block, + paletteInfo); } return this.selectedBlockCost; @@ -686,7 +845,8 @@ internal static partial class Av1IntraSuperblockEncoder int lumaArea, int chromaArea, Av1MacroBlockModeInfo modeInfo, - Av1EncoderBlockStruct block) + Av1EncoderBlockStruct block, + Av1EncoderPaletteInfo paletteInfo) { Av1BlockSize blockSize = modeInfo.Block.BlockSize; Av1TransformSize transformSize = modeInfo.Block.TransformSize; @@ -707,21 +867,27 @@ internal static partial class Av1IntraSuperblockEncoder Span lumaStates = this.coefficientBuffer.GetTransformBlockSpan(this.superblock.Index, Av1Plane.Y); - Av1EncoderTransformBlockState lumaState = - lumaStates[lumaArea / Av1EncoderCoefficientBuffer.TransformBlockUnitCoefficientCount]; - Span lumaCoefficients = this.coefficientBuffer.GetPlaneSpan(this.superblock.Index, Av1Plane.Y); - byte lumaContext = Av1SymbolContextHelper.GetCoefficientContext( - lumaCoefficients[lumaArea..], + PublishCoefficientContexts( + this.picture.LuminanceDcSignLevelCoefficientNeighbors[tileIndex], + blockOrigin, + blockSize, transformSize, - lumaState.TransformType, - lumaState.EndOfBlock); + lumaCoefficients[lumaArea..], + lumaStates[(lumaArea / Av1EncoderCoefficientBuffer.TransformBlockUnitCoefficientCount)..]); - this.picture.LuminanceDcSignLevelCoefficientNeighbors[tileIndex].UnitModeWrite( - lumaContext, - blockOrigin, - blockDimensions, - Av1NeighborArrayUnit.UnitMask.Top | Av1NeighborArrayUnit.UnitMask.Left); + if (this.picture.Parent.FrameHeader.AllowScreenContentTools) + { + const Av1NeighborArrayUnit.UnitMask PaletteContextMask = + Av1NeighborArrayUnit.UnitMask.Top | + Av1NeighborArrayUnit.UnitMask.Left; + + this.picture.PaletteContexts[tileIndex].UnitModeWrite( + paletteInfo, + blockOrigin, + blockDimensions, + PaletteContextMask); + } if (!block.HasChroma) { @@ -753,56 +919,111 @@ internal static partial class Av1IntraSuperblockEncoder Span redStates = this.coefficientBuffer.GetTransformBlockSpan(this.superblock.Index, Av1Plane.V); - Av1EncoderTransformBlockState blueState = blueStates[chromaStateIndex]; - Av1EncoderTransformBlockState redState = redStates[chromaStateIndex]; Span blueCoefficients = this.coefficientBuffer.GetPlaneSpan(this.superblock.Index, Av1Plane.U); Span redCoefficients = this.coefficientBuffer.GetPlaneSpan(this.superblock.Index, Av1Plane.V); - byte blueContext = Av1SymbolContextHelper.GetCoefficientContext( - blueCoefficients[chromaArea..], + PublishCoefficientContexts( + this.picture.CbDcSignLevelCoefficientNeighbors[tileIndex], + chromaOrigin, + chromaBlockSize, chromaTransformSize, - blueState.TransformType, - blueState.EndOfBlock); + blueCoefficients[chromaArea..], + blueStates[chromaStateIndex..]); - byte redContext = Av1SymbolContextHelper.GetCoefficientContext( - redCoefficients[chromaArea..], + PublishCoefficientContexts( + this.picture.CrDcSignLevelCoefficientNeighbors[tileIndex], + chromaOrigin, + chromaBlockSize, chromaTransformSize, - redState.TransformType, - redState.EndOfBlock); + redCoefficients[chromaArea..], + redStates[chromaStateIndex..]); + } - Size chromaDimensions = new(chromaBlockSize.GetWidth(), chromaBlockSize.GetHeight()); - this.picture.CbDcSignLevelCoefficientNeighbors[tileIndex].UnitModeWrite( - blueContext, - chromaOrigin, - chromaDimensions, - Av1NeighborArrayUnit.UnitMask.Top | Av1NeighborArrayUnit.UnitMask.Left); + private static void PublishCoefficientContexts( + Av1NeighborArrayUnit neighbors, + Point blockOrigin, + Av1BlockSize blockSize, + Av1TransformSize transformSize, + ReadOnlySpan coefficients, + ReadOnlySpan states) + { + const Av1NeighborArrayUnit.UnitMask EdgeMask = + Av1NeighborArrayUnit.UnitMask.Top | + Av1NeighborArrayUnit.UnitMask.Left; - this.picture.CrDcSignLevelCoefficientNeighbors[tileIndex].UnitModeWrite( - redContext, - chromaOrigin, - chromaDimensions, - Av1NeighborArrayUnit.UnitMask.Top | Av1NeighborArrayUnit.UnitMask.Left); + int blockWidth = blockSize.GetWidth(); + int blockHeight = blockSize.GetHeight(); + int transformWidth = transformSize.GetWidth(); + int transformHeight = transformSize.GetHeight(); + int transformSampleCount = transformSize.GetSize2d(); + int transformIndex = 0; + int coefficientOffset = 0; + + // Uniform transform blocks are retained and written in raster order. Publishing that same tiling + // preserves the distinct top and left contexts consumed by the next coding block in a dry run. + for (int row = 0; row < blockHeight; row += transformHeight) + { + for (int column = 0; column < blockWidth; column += transformWidth) + { + Av1EncoderTransformBlockState state = states[transformIndex++]; + byte context = Av1SymbolContextHelper.GetCoefficientContext( + coefficients[coefficientOffset..], + transformSize, + state.TransformType, + state.EndOfBlock); + + neighbors.UnitModeWrite( + context, + blockOrigin + new Size(column, row), + new Size(transformWidth, transformHeight), + EdgeMask); + + coefficientOffset += transformSampleCount; + } + } } - private void SavePartitionTrialContexts(Point blockOrigin, ushort tileIndex) + private void SavePartitionTrialContexts( + Point blockOrigin, + ushort tileIndex, + Av1BlockSize blockSize) { - Span storage = this.blockWorkspace.PartitionContexts; + Span storage = this.blockWorkspace.GetPartitionContextStorage(blockSize); int offset = 0; + SaveNeighborEdges( + this.picture.PartitionContexts[tileIndex], + blockOrigin, + blockSize.Get4x4WideCount(), + blockSize.Get4x4HighCount(), + storage, + ref offset); + SaveNeighborEdges( this.picture.LuminanceDcSignLevelCoefficientNeighbors[tileIndex], blockOrigin, - Av1BlockSize.Block8x8.Get4x4WideCount(), - Av1BlockSize.Block8x8.Get4x4HighCount(), + blockSize.Get4x4WideCount(), + blockSize.Get4x4HighCount(), storage, ref offset); SaveNeighborEdges( this.picture.TransformFunctionContexts[tileIndex], blockOrigin, - Av1BlockSize.Block8x8.Get4x4WideCount(), - Av1BlockSize.Block8x8.Get4x4HighCount(), + blockSize.Get4x4WideCount(), + blockSize.Get4x4HighCount(), storage, ref offset); + if (this.picture.Parent.FrameHeader.AllowScreenContentTools) + { + SaveNeighborEdges( + this.picture.PaletteContexts[tileIndex], + blockOrigin, + blockSize.Get4x4WideCount(), + blockSize.Get4x4HighCount(), + storage, + ref offset); + } + if (this.source.IsMonochrome) { return; @@ -816,7 +1037,7 @@ internal static partial class Av1IntraSuperblockEncoder subsamplingX, subsamplingY); - Av1BlockSize chromaBlockSize = Av1BlockSize.Block8x8.GetSubsampled( + Av1BlockSize chromaBlockSize = blockSize.GetSubsampled( colorConfig.SubSamplingX, colorConfig.SubSamplingY); @@ -837,26 +1058,48 @@ internal static partial class Av1IntraSuperblockEncoder ref offset); } - private void RestorePartitionTrialContexts(Point blockOrigin, ushort tileIndex) + private void RestorePartitionTrialContexts( + Point blockOrigin, + ushort tileIndex, + Av1BlockSize blockSize) { - ReadOnlySpan storage = this.blockWorkspace.PartitionContexts; + ReadOnlySpan storage = this.blockWorkspace.GetPartitionContextStorage(blockSize); int offset = 0; + RestoreNeighborEdges( + this.picture.PartitionContexts[tileIndex], + blockOrigin, + blockSize.Get4x4WideCount(), + blockSize.Get4x4HighCount(), + storage, + ref offset); + RestoreNeighborEdges( this.picture.LuminanceDcSignLevelCoefficientNeighbors[tileIndex], blockOrigin, - Av1BlockSize.Block8x8.Get4x4WideCount(), - Av1BlockSize.Block8x8.Get4x4HighCount(), + blockSize.Get4x4WideCount(), + blockSize.Get4x4HighCount(), storage, ref offset); RestoreNeighborEdges( this.picture.TransformFunctionContexts[tileIndex], blockOrigin, - Av1BlockSize.Block8x8.Get4x4WideCount(), - Av1BlockSize.Block8x8.Get4x4HighCount(), + blockSize.Get4x4WideCount(), + blockSize.Get4x4HighCount(), storage, ref offset); + if (this.picture.Parent.FrameHeader.AllowScreenContentTools) + { + RestoreNeighborEdges( + this.picture.PaletteContexts[tileIndex], + blockOrigin, + blockSize.Get4x4WideCount(), + blockSize.Get4x4HighCount(), + storage, + ref offset); + } + if (this.source.IsMonochrome) { return; @@ -870,7 +1113,7 @@ internal static partial class Av1IntraSuperblockEncoder subsamplingX, subsamplingY); - Av1BlockSize chromaBlockSize = Av1BlockSize.Block8x8.GetSubsampled( + Av1BlockSize chromaBlockSize = blockSize.GetSubsampled( colorConfig.SubSamplingX, colorConfig.SubSamplingY); @@ -891,32 +1134,46 @@ internal static partial class Av1IntraSuperblockEncoder ref offset); } - private static void SaveNeighborEdges( - Av1NeighborArrayUnit neighbors, + private static void SaveNeighborEdges( + Av1NeighborArrayUnit neighbors, Point blockOrigin, int width, int height, Span storage, ref int offset) + where T : struct { - neighbors.Top.Slice(neighbors.GetTopIndex(blockOrigin), width).CopyTo(storage[offset..]); - offset += width; - neighbors.Left.Slice(neighbors.GetLeftIndex(blockOrigin), height).CopyTo(storage[offset..]); - offset += height; + Span top = MemoryMarshal.AsBytes( + neighbors.Top.Slice(neighbors.GetTopIndex(blockOrigin), width)); + + top.CopyTo(storage[offset..]); + offset += top.Length; + Span left = MemoryMarshal.AsBytes( + neighbors.Left.Slice(neighbors.GetLeftIndex(blockOrigin), height)); + + left.CopyTo(storage[offset..]); + offset += left.Length; } - private static void RestoreNeighborEdges( - Av1NeighborArrayUnit neighbors, + private static void RestoreNeighborEdges( + Av1NeighborArrayUnit neighbors, Point blockOrigin, int width, int height, ReadOnlySpan storage, ref int offset) + where T : struct { - storage.Slice(offset, width).CopyTo(neighbors.Top[neighbors.GetTopIndex(blockOrigin)..]); - offset += width; - storage.Slice(offset, height).CopyTo(neighbors.Left[neighbors.GetLeftIndex(blockOrigin)..]); - offset += height; + Span top = MemoryMarshal.AsBytes( + neighbors.Top.Slice(neighbors.GetTopIndex(blockOrigin), width)); + + storage.Slice(offset, top.Length).CopyTo(top); + offset += top.Length; + Span left = MemoryMarshal.AsBytes( + neighbors.Left.Slice(neighbors.GetLeftIndex(blockOrigin), height)); + + storage.Slice(offset, left.Length).CopyTo(left); + offset += left.Length; } private long GetRegularBlockCost( diff --git a/src/ImageSharp/Formats/Heif/Av1/Tiling/Av1PartitionContext.cs b/src/ImageSharp/Formats/Heif/Av1/Tiling/Av1PartitionContext.cs index 31af3be2d8..ef7e9b738a 100644 --- a/src/ImageSharp/Formats/Heif/Av1/Tiling/Av1PartitionContext.cs +++ b/src/ImageSharp/Formats/Heif/Av1/Tiling/Av1PartitionContext.cs @@ -2,6 +2,7 @@ // Licensed under the Six Labors Split License. using System.Numerics; +using System.Runtime.InteropServices; namespace SixLabors.ImageSharp.Formats.Heif.Av1.Tiling; @@ -12,8 +13,14 @@ namespace SixLabors.ImageSharp.Formats.Heif.Av1.Tiling; /// Each set bit records a split at one block-size level. For example, 11111 records splits from /// 128 by 128 through 8 by 8, while 10000 records only the 128 by 128 split. /// +[StructLayout(LayoutKind.Sequential, Size = StorageSize)] internal struct Av1PartitionContext : IMinMaxValue { + /// + /// The packed size of the above and left context bytes. + /// + public const int StorageSize = 2; + /// /// Maps each block size to the five-bit context stored for an above neighbor. /// diff --git a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs index c26d3a4f78..6f3cbc3012 100644 --- a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs +++ b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1EncoderFrameTests.cs @@ -216,6 +216,64 @@ public class Av1EncoderFrameTests Assert.Equal(new Size(Size, Size), decoded.Size); } + [Fact] + public void EncodeEffortNineSelectsSixteenBySixteenVerticalPartition() + { + const int Size = 32; + using Image source = new(Size, Size); + for (int y = 0; y < Size; y++) + { + Span row = source.Frames.RootFrame.PixelBuffer.DangerousGetRowSpan(y); + for (int x = 0; x < Size; x++) + { + byte value = 128; + if (x == 15 && y >= 16) + { + value = (byte)(24 + ((y - 16) * 13)); + } + else if (y == 15 && x >= 16) + { + value = (byte)(16 + ((x - 16) * 15)); + } + else if (x >= 16 && y >= 16) + { + // The left 8x16 half repeats its external left edge, while the right half repeats + // its external top edge. One 16x16 predictor cannot reproduce both surfaces. + value = x < 24 + ? (byte)(24 + ((y - 16) * 13)) + : (byte)(16 + ((x - 16) * 15)); + } + + row[x] = new Rgba32(value, value, value); + } + } + + using MemoryStream stream = new(); + _ = Av1FrameEncoder.Encode( + Configuration.Default, + source.Frames.RootFrame, + stream, + CreateColorConfig(Av1BitDepth.EightBit, Av1ColorFormat.Yuv400), + qIndex: 4, + effort: 9); + + byte[] payload = stream.ToArray(); + using Av1Decoder decoder = new(Configuration.Default); + using Image decoded = decoder.Decode(payload); + Av1FrameInfo frameInfo = Assert.IsType(decoder.FrameInfo); + for (int modeInfoY = 4; modeInfoY < 8; modeInfoY++) + { + for (int modeInfoX = 4; modeInfoX < 8; modeInfoX++) + { + Assert.Equal( + Av1BlockSize.Block8x16, + frameInfo.GetModeInfoAt(new Point(modeInfoX, modeInfoY)).BlockSize); + } + } + + Assert.Equal(new Size(Size, Size), decoded.Size); + } + [Theory] [InlineData(EightBit)] [InlineData(TenBit)] diff --git a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs index 3aaa4e4835..7b2effd840 100644 --- a/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs +++ b/tests/ImageSharp.Tests/Formats/Heif/Av1/Av1TransformBlockEncoderTests.cs @@ -617,10 +617,17 @@ public class Av1TransformBlockEncoderTests Av1EncoderIntraBlockCopyWorkspace intraBlockCopyWorkspace = workspace.GetIntraBlockCopyWorkspace(); - Assert.Equal(17, modeWorkspace.GetReferenceSamples(3).Length); + Assert.Equal( + (2 * Av1EncoderModeDecisionWorkspace.MaximumBlockDimension) + 1, + modeWorkspace.GetReferenceSamples(3).Length); + Assert.Equal(Av1EncoderModeDecisionWorkspace.MaximumSampleCount, modeWorkspace.GetCandidateReconstruction(1).Length); Assert.Equal(Av1EncoderModeDecisionWorkspace.MaximumSampleCount, modeWorkspace.GetCandidateCoefficients(1).Length); - Assert.Equal(Av1ChromaFromLumaContext.BufferLine * 8, modeWorkspace.ChromaFromLumaSamples.Length); + Assert.Equal( + Av1ChromaFromLumaContext.BufferLine * + Av1EncoderModeDecisionWorkspace.MaximumBlockDimension, + modeWorkspace.ChromaFromLumaSamples.Length); + Assert.Equal(Av1ChromaFromLumaMath.AlphaCandidateCount, modeWorkspace.GetChromaFromLumaRates(1).Length); Assert.Equal(Av1ChromaFromLumaMath.AlphaCandidateCount, modeWorkspace.GetChromaFromLumaDistortions(1).Length); Assert.Equal(Av1EncoderModeDecisionWorkspace.MaximumSampleCount, paletteWorkspace.GetPrediction(1).Length);