libyuv

mirror of https://chromium.googlesource.com/libyuv/libyuv synced 2025-12-08 01:36:47 +08:00

Author	SHA1	Message	Date
George Steed	15f2ae7d70	[AArch64] Add SME impls of ScaleARGBRowDown2{,Linear,Box} Mostly just straightforward copies of the Neon code ported to Streaming-SVE, these follow the same pattern as the prior ScaleRowDown2 and ScaleUVRowDown2 SME kernels, but operating on 32-bit ARGB tuples rather than 8-bit data or 16-bit UV tuples. These is no benefit from this kernel when the SVE vector length is only 128 bits, so skip writing a non-streaming SVE implementation. Change-Id: I15600c2498cc592f5ea1d97b78fafec327de7947 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6070783 Reviewed-by: Frank Barchard <fbarchard@chromium.org> Reviewed-by: Justin Green <greenjustin@google.com>	2024-12-12 01:19:20 -08:00
George Steed	7391559cb4	[AArch64] Add SME implementation of MergeUVRow{,_16} Mostly just a straightforward copy of the Neon code ported to Streaming-SVE, we can use predication to avoid needing an `Any` kernel and use ST2 to avoid needing a separate ZIP instruction. These is no benefit from this kernel when the SVE vector length is only 128 bits, so skip writing a non-streaming SVE implementation. Change-Id: I5ae36afe699b88f119dc545e49c59c5d85e98742 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6070785 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-12-12 01:16:19 -08:00
George Steed	8f659daffd	[AArch64] Add SVE2 implementations of NV{12,21}ToRGB24Row Now that we have the `_2X` versions of the macros we can use these to implement `ToRGB24` kernels. These cannot use the bottom/top approach previously used by other SVE kernels since there are three rather than two or four elements each. Reduction in runtimes observed compared to the existing Neon implementations: \| NV12ToRGB24Row \| NV21ToRGB24Row Cortex-A510 \| -60.7% \| -60.7% Cortex-A520 \| -46.0% \| -46.0% Cortex-A715 \| -25.2% \| -25.2% Cortex-A720 \| -25.2% \| -25.2% Cortex-X2 \| -28.9% \| -29.0% Cortex-X3 \| -28.2% \| -28.1% Cortex-X4 \| -30.8% \| -30.7% Cortex-X925 \| -28.8% \| -28.9% Change-Id: I39853d124bfdcac38584109870b398b8ecd5b632 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6067149 Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-12-04 17:51:08 +00:00
George Steed	9144583f22	[AArch64] Add SME impls of MultiplyRow_16 and ARGBMultiplyRow Mostly just a translation of the existing Neon code to SME. Change-Id: Ic3d6b8ac774c9a1bb9204ed6c78c8802668bffe9 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6067147 Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-12-03 22:11:19 +00:00
George Steed	9a9752134e	[AArch64] Add Neon implementation of ScaleRowDown2Linear_16 Reduction in runtime observed relative to the auto-vectorized C implementation compiled with LLVM 19: Cortex-A55: -13.7% Cortex-A510: -49.0% Cortex-A520: -32.0% Cortex-A76: -34.3% Cortex-A710: -56.7% Cortex-A715: -45.4% Cortex-A720: -44.7% Cortex-X1: -70.6% Cortex-X2: -67.9% Cortex-X3: -72.2% Cortex-X4: -40.0% Cortex-X925: -24.1% Bug: b/42280942 Change-Id: I977899a2239e752400c9901f4d8482a76841269a Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6040154 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-25 21:10:26 +00:00
George Steed	11c57f4f12	[AArch64] Add Neon implementation of ScaleRowDown2_16_NEON The auto-vectorized implementation unrolls to process 32 elements per iteration, so unroll the new Neon implementation to match and avoid a performance regression on little cores. Performance relative to the auto-vectorized C implementation compiled with LLVM 19: Cortex-A55: -35.8% Cortex-A510: -20.4% Cortex-A520: -22.1% Cortex-A76: -54.8% Cortex-A710: -44.5% Cortex-A715: -31.1% Cortex-A720: -31.4% Cortex-X1: -48.5% Cortex-X2: -47.8% Cortex-X3: -47.6% Cortex-X4: -51.1% Cortex-X925: -14.6% Bug: b/42280942 Change-Id: Ib4e89ba230d554f2717052e934ca0e8a109ccc42 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6040153 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-25 21:10:05 +00:00
George Steed	952d6a282f	[AArch64] Enable use of ScaleRowDown2Box_16_NEON The #ifdef surrounding the use of this kernel is never defined and ScaleRowDown2_16_NEON does not exist, so add the missing #define and remove the use of ScaleRowDown2_16_NEON for now. Additionally since there is no implementation of this kernel for 32-bit Arm, restrict the define to only be present on AArch64. Bug: b/42280942 Change-Id: Icc35c145c1bad1c0df2933a2d8bc7dcf7fe63cb7 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6040152 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-24 19:58:00 +00:00
George Steed	9ed07258c7	[AArch64] Add SVE2 implementation of I410ToAR30Row Observed reduction in runtime compared to the existing Neon code: Cortex-A510: -18.1% Cortex-A520: -6.0% Cortex-A715: -22.0% Cortex-A720: -21.1% Cortex-X2: -9.4% Cortex-X3: -12.0% Cortex-X4: -7.6% Cortex-X925: -5.8% Bug: b/42280942 Change-Id: I853a028e08f1f1076ac20cd9c7f4f8ac8a211ac1 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6023584 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-23 00:59:55 +00:00
George Steed	3dd047733e	[AArch64] Add SVE2 implementation of I410AlphaToARGBRow Observed reduction in runtime compared to the existing Neon code: Cortex-A510: -37.2% Cortex-A520: -6.9% Cortex-A715: -14.8% Cortex-A720: -16.0% Cortex-X2: -14.8% Cortex-X3: -17.5% Cortex-X4: -12.8% Cortex-X925: -13.0% Bug: b/42280942 Change-Id: I1977fd1e1dfac25021724483fd89c6ff3e227d8b Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6023582 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-23 00:58:11 +00:00
George Steed	e84d809348	[AArch64] Add SVE2 implementation of I410ToARGBRow Observed reduction in runtime compared to the existing Neon code: Cortex-A510: -37.9% Cortex-A520: -9.2% Cortex-A715: -14.3% Cortex-A720: -14.2% Cortex-X2: -10.9% Cortex-X3: -11.1% Cortex-X4: -12.5% Cortex-X925: -10.6% Bug: b/42280942 Change-Id: I6720b07c900c7dfbd849ee38e413e98b9374dac2 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6023581 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-23 00:54:48 +00:00
George Steed	7c9c72ab4b	[AArch64] Add SVE2 implementation of I210ToAR30Row Observed reduction in runtime compared to the existing Neon code: Cortex-A510: -15.5% Cortex-A520: -3.8% Cortex-A715: -15.8% Cortex-A720: -15.8% Cortex-X2: -7.9% Cortex-X3: -6.5% Cortex-X4: -5.0% Cortex-X925: -5.3% Bug: b/42280942 Change-Id: I5171537fd125b3214d25a0ae503a8f40dbeb6042 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6023583 Reviewed-by: Frank Barchard <fbarchard@chromium.org> Reviewed-by: Justin Green <greenjustin@google.com>	2024-11-23 00:53:16 +00:00
George Steed	fc3569ad27	[AArch64] Add SVE2 implementation of I210AlphaToARGBRow Observed reduction in runtime compared to the existing Neon code: Cortex-A510: -33.9% Cortex-A520: -4.2% Cortex-A715: -22.0% Cortex-A720: -22.4% Cortex-X2: -14.6% Cortex-X3: -14.5% Cortex-X4: -11.6% Cortex-X925: -12.6% Bug: b/42280942 Change-Id: Ifb4ed7a865c369d584af498cc65b84d065cfb207 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6023580 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-23 00:47:32 +00:00
George Steed	50108f29fb	[AArch64] Add SVE2 implementation of I212ToAR30Row Observed reduction in runtime compared to the existing Neon code: Cortex-A510: -15.4% Cortex-A520: -3.8% Cortex-A715: -15.7% Cortex-A720: -15.6% Cortex-X2: -7.9% Cortex-X3: -5.7% Cortex-X4: -5.3% Cortex-X925: -4.8% Bug: b/42280942 Change-Id: I99846820682687c8e0f52d05f5aa3d50369fe0a2 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6025829 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-23 00:27:57 +00:00
George Steed	305a7a4ede	[AArch64] Add SVE2 implementation of I212ToARGBRow Observed reduction in runtime compared to the existing Neon code: Cortex-A510: -34.5% Cortex-A520: -6.5% Cortex-A715: -10.1% Cortex-A720: -16.1% Cortex-X2: -11.9% Cortex-X3: -11.9% Cortex-X4: -9.3% Cortex-X925: -11.2% Bug: b/42280942 Change-Id: Idc30e69552f7d227217ac7011a786210b11e4752 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6025828 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-23 00:21:27 +00:00
Frank Barchard	595146434a	HalfFloat fix SigIll on aarch64 - Remove special case Scale of 1 which used fp16 cvt but requires cpuid - Port aarch64 to aarch32 - Use C for aarch32 with small (denormal) scale value Bug: 377693555 Change-Id: I38e207e79ac54907ed6e65118b8109288fddb207 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6043392 Reviewed-by: Wan-Teh Chang <wtc@google.com>	2024-11-22 22:08:00 +00:00
Frank Barchard	307b951229	Add CopyPlane_Unaligned, _Any and _Invert tests/benchmarksCpuId test - Add AMD_ERMSB detect for ERMS on AMD Bug: 379457420 Change-Id: I608568556024faf19abe4d0662aeeee553a0a349 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6032852 Reviewed-by: Wan-Teh Chang <wtc@google.com>	2024-11-19 23:53:05 +00:00
Frank Barchard	1c501a8f3f	CpuId test FSMR - Fast Short Rep Movsb - Renumber cpuid bits to use low byte to ID the type of CPU and upper 24 bits for features Intel CPUs starting at Icelake support FSMR adl:Has FSMR 0x8000 arl:Has FSMR 0x0 bdw:Has FSMR 0x0 clx:Has FSMR 0x0 cnl:Has FSMR 0x0 cpx:Has FSMR 0x0 emr:Has FSMR 0x8000 glm:Has FSMR 0x0 glp:Has FSMR 0x0 gnr:Has FSMR 0x8000 gnr256:Has FSMR 0x8000 hsw:Has FSMR 0x0 icl:Has FSMR 0x8000 icx:Has FSMR 0x8000 ivb:Has FSMR 0x0 knl:Has FSMR 0x0 knm:Has FSMR 0x0 lnl:Has FSMR 0x8000 mrm:Has FSMR 0x0 mtl:Has FSMR 0x8000 nhm:Has FSMR 0x0 pnr:Has FSMR 0x0 rpl:Has FSMR 0x8000 skl:Has FSMR 0x0 skx:Has FSMR 0x0 slm:Has FSMR 0x0 slt:Has FSMR 0x0 snb:Has FSMR 0x0 snr:Has FSMR 0x0 spr:Has FSMR 0x8000 srf:Has FSMR 0x0 tgl:Has FSMR 0x8000 tnt:Has FSMR 0x0 wsm:Has FSMR 0x0 Intel CPUs starting at Ivybridge support ERMS adl:Has ERMS 0x4000 arl:Has ERMS 0x4000 bdw:Has ERMS 0x4000 clx:Has ERMS 0x4000 cnl:Has ERMS 0x4000 cpx:Has ERMS 0x4000 emr:Has ERMS 0x4000 glm:Has ERMS 0x4000 glp:Has ERMS 0x4000 gnr:Has ERMS 0x4000 gnr256:Has ERMS 0x4000 hsw:Has ERMS 0x4000 icl:Has ERMS 0x4000 icx:Has ERMS 0x4000 ivb:Has ERMS 0x4000 knl:Has ERMS 0x4000 knm:Has ERMS 0x4000 lnl:Has ERMS 0x4000 mrm:Has ERMS 0x0 mtl:Has ERMS 0x4000 nhm:Has ERMS 0x0 pnr:Has ERMS 0x0 rpl:Has ERMS 0x4000 skl:Has ERMS 0x4000 skx:Has ERMS 0x4000 slm:Has ERMS 0x4000 slt:Has ERMS 0x0 snb:Has ERMS 0x0 snr:Has ERMS 0x4000 spr:Has ERMS 0x4000 srf:Has ERMS 0x4000 tgl:Has ERMS 0x4000 tnt:Has ERMS 0x4000 wsm:Has ERMS 0x0 Change-Id: I18e5a3905f2691ab66d4d0cb6f668c0a0ff72d37 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6027541 Reviewed-by: richard winterton <rrwinterton@gmail.com>	2024-11-18 17:56:45 +00:00
Frank Barchard	75f7cfdde5	SplitRGB for SSE4 and AVX2 libyuv_test '--gunit_filter=SplitRGB' --libyuv_width=640 --libyuv_height=360 --libyuv_repeat=100000 --libyuv_flags=-1 --libyuv_cpu_info=-1 Note: Google Test filter = SplitRGB Skylake Xeon x86 32 bit AVX2 LibYUVPlanarTest.SplitRGBPlane_Opt (4143 ms) SSE4 LibYUVPlanarTest.SplitRGBPlane_Opt (4543 ms) SSSE3 LibYUVPlanarTest.SplitRGBPlane_Opt (5346 ms) C LibYUVPlanarTest.SplitRGBPlane_Opt (22965 ms) Skylake Xeon x86 64 bit AVX2 LibYUVPlanarTest.SplitRGBPlane_Opt (4470 ms) SSE4 LibYUVPlanarTest.SplitRGBPlane_Opt (4723 ms) SSSE3 LibYUVPlanarTest.SplitRGBPlane_Opt (5465 ms) C LibYUVPlanarTest.SplitRGBPlane_Opt (4707 ms) Bug: 379186682 Change-Id: Idce67a4ded836f2ee31854aa06f3903e7bcb7791 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6024314 Reviewed-by: richard winterton <rrwinterton@gmail.com>	2024-11-15 00:46:25 +00:00
George Steed	823d960afc	[AArch64] Add SVE2 implementations of {P210,P410}ToAR30Row Observed reductions in runtime compared to the existing Neon code: \| P210ToAR30Row \| P410ToAR30Row Cortex-A510 \| -16.5% \| -21.2% Cortex-A520 \| (!) +2.7% \| -8.7% Cortex-A715 \| -6.1% \| -6.1% Cortex-A720 \| -6.2% \| -5.9% Cortex-X2 \| -4.1% \| -4.2% Cortex-X3 \| -4.2% \| -4.2% Cortex-X4 \| -1.2% \| -1.2% Cortex-X925 \| -3.6% \| -2.8% Bug: b/42280942 Change-Id: I40723a370fad1ccb53f8ccd9d32cddb502500dd6 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6023036 Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-14 16:52:21 +00:00
George Steed	0ddf3f7b90	[AArch64] Add SVE2 implementation of I210ToARGBRow Observed reduction in runtime compared to the existing Neon code: Cortex-A510: -34.5% Cortex-A520: -6.5% Cortex-A715: -10.1% Cortex-A720: -13.9% Cortex-X2: -11.9% Cortex-X3: -11.6% Cortex-X4: -9.5% Cortex-X925: -11.5% Bug: b/42280942 Change-Id: Ie97dc3b5efd021ecfea14d4c477cc205191e09c3 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/6023037 Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-14 16:36:41 +00:00
George Steed	5b906a0ec8	[AArch64] Add SVE2 implementation of P410ToARGBRow Observed reduction in runtime compared to the existing Neon code: Cortex-A510: -34.7% Cortex-A520: -2.4% Cortex-A715: -18.7% Cortex-A720: -18.8% Cortex-X2: -7.7% Cortex-X3: -8.9% Cortex-X4: +1.0% (!) Cortex-X925: -8.3% Bug: b/42280942 Change-Id: I90dca0573887a9a24e2172378a9e0eb6812e2131 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5975321 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-12 18:34:56 +00:00
George Steed	b753822d47	[AArch64] Add SVE2 implementation of P210ToARGBRow Observed reduction in runtime compared to the existing Neon code: Cortex-A510: -32.8% Cortex-A520: +8.7% (!) Cortex-A715: -18.9% Cortex-A720: -18.9% Cortex-X2: -7.9% Cortex-X3: -8.8% Cortex-X4: +1.0% (!) Cortex-X925: -8.6% Bug: b/42280942 Change-Id: Ibe557500c3788b4fb39372c92b2f42ba216e6fea Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5975320 Reviewed-by: Frank Barchard <fbarchard@chromium.org> Reviewed-by: Justin Green <greenjustin@google.com>	2024-11-12 18:32:55 +00:00
George Steed	721ad4aa18	[AArch64] Add SME implementation of ScaleUVRowDown2Box There is no benefit from an SVE version of this kernel for devices with an SVE vector length of 128-bits, so skip directly to SME instead. We do not use the ZA tile here, so this is a purely streaming-SVE (SSVE) implementation. Change-Id: Ie15bb4e7484b61e78f405ad4e8a8a7bbb66b7edb Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5979727 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-12 18:30:30 +00:00
George Steed	576218dbce	[AArch64] Add SME implementation of ScaleUVRowDown2Linear There is no benefit from an SVE version of this kernel for devices with an SVE vector length of 128-bits, so skip directly to SME instead. We do not use the ZA tile here, so this is a purely streaming-SVE (SSVE) implementation. Change-Id: I401eb6ad14b3159917c8e3a79ab20dde318d28b6 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5979726 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-12 18:28:57 +00:00
George Steed	551cee7845	[AArch64] Add SME implementation of ScaleUVRowDown2 There is no benefit from an SVE version of this kernel for devices with an SVE vector length of 128-bits, so skip directly to SME instead. We do not use the ZA tile here, so this is a purely streaming-SVE (SSVE) implementation. Change-Id: Ic4ba5f97dc57afc558c08a57e9b5009d6e487e0f Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5979725 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-12 18:24:28 +00:00
George Steed	5c12e0b2de	[AArch64] Add SVE2 implementations of HalfFloat{,1}Row For HalfFloat1Row, SVE has direct 16-bit integer to half-float conversion instructions so there is no need to widen to 32-bits. For HalfFloatRow, SVE zero-extending loads avoid the need for seperate UXTL(2) instructions. Observed reductions in runtime compared to the existing Neon code: \| HalfFloat1Row \| HalfFloatRow Cortex-A510 \| -38.3% \| -17.3% Cortex-A520 \| -37.6% \| -18.8% Cortex-A720 \| -50.1% \| -7.8% Cortex-X2 \| -50.2% \| -0.4% Cortex-X4 \| -51.5% \| -12.5% Bug: b/42280942 Change-Id: I445071ccd453113144ce42d465ba03c9ee89ec9e Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5975319 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-07 18:53:00 +00:00
George Steed	f27b983f38	[AArch64] Add SVE2 implementation of DivideRow_16 SVE contains the UMULH instruction which allows us to multiply and take the high half of the result in a single instruction rather than needing separate widening multiply and then narrowing shift steps. Observed reduction in runtime compared to the existing Neon code: Cortex-A510: -21.2% Cortex-A520: -20.9% Cortex-A715: -47.9% Cortex-A720: -47.6% Cortex-X2: -5.2% Cortex-X3: -2.6% Cortex-X4: -32.4% Cortex-X925: -1.5% Bug: b/42280942 Change-Id: I25154699b17772db1fb5cb84c049919181d86f4b Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5975318 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-07 18:46:02 +00:00
George Steed	aec4b4e22e	[AArch64] Add SME implementation of ScaleRowDown2Box There is no benefit from an SVE version of this kernel for devices with an SVE vector length of 128-bits, so skip directly to SME instead. We do not use the ZA tile here, so this is a purely streaming-SVE (SSVE) implementation. Change-Id: I5021aeda30f4c5f1aa4cc6326c8d7886851d2c09 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5913885 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-11-07 18:42:21 +00:00
George Steed	51d07554a0	[AArch64] Add SME implementation of ScaleRowDown2Linear There is no benefit from an SVE version of this kernel for devices with an SVE vector length of 128-bits, so skip directly to SME instead. We do not use the ZA tile here, so this is a purely streaming-SVE (SSVE) implementation. Change-Id: Ie6b91bd4407130ba2653838088e81e72e4460f68 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5913884 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-10-30 17:57:15 +00:00
George Steed	593965cea2	[AArch64] Add SME implementation of ScaleRowDown2 Including associated changes for adding a new scale_sme.cc file. There is no benefit from an SVE version of this kernel for devices with an SVE vector length of 128-bits, so skip directly to SME instead. We do not use the ZA tile here, so this is a purely streaming-SVE (SSVE) implementation. Change-Id: I47d149613fbabd8c203605a809811f1a668e8fb7 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5913883 Reviewed-by: Frank Barchard <fbarchard@chromium.org> Reviewed-by: Justin Green <greenjustin@google.com>	2024-10-30 17:56:41 +00:00
George Steed	237f39cb8c	[AArch64] Add SME implementation of I444ToARGBRow This is based on an unrolled version of the existing SVE2 code. The implementation in this case is a pure streaming-SVE (SSVE) implementation based on the existing SVE2 implementation, we do not use the ZA tile. Change-Id: I83d8e58aafd814125b3446fb1c9ec4a5fb56fe3e Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5913882 Reviewed-by: Frank Barchard <fbarchard@chromium.org> Reviewed-by: Justin Green <greenjustin@google.com>	2024-10-29 18:10:23 +00:00
George Steed	22c5c18778	[AArch64] Add SME implementation of I422ToARGBRow Including addition of a new row_sme.cc file and associated infrastructure. The actual implementation in this case is a pure streaming-SVE (SSVE) implementation based on the existing SVE2 implementation, we do not use the ZA tile. Change-Id: Ibc132c55de8d41a107e563b95f842323fef94444 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5913881 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-10-29 05:49:28 +00:00
George Steed	22ac86800e	[AArch64] Add SVE2 implementation of I422ToARGB4444Row This makes use of the same approach as the Neon code to avoid redundant narrowing and then widening shifts by instead placing the values at the top portion of the lanes and then shifting down from there instead. Observed reduction in runtime compared to the existing Neon code: Cortex-A510: -35.5% Cortex-A520: -38.2% Cortex-A715: -19.8% Cortex-A720: -19.8% Cortex-X2: -24.2% Cortex-X3: -24.1% Cortex-X4: -21.6% Cortex-X925: -19.5% Bug: b/42280942 Change-Id: I0a916600e7bdee0f5480ea843b44ab046bb3d082 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5802968 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-10-24 21:27:39 +00:00
George Steed	f4eaeca22a	[AArch64] Add SVE2 implementation of I422ToARGB1555Row This makes use of the same approach as the Neon code to avoid redundant narrowing and then widening shifts by instead placing the values at the top portion of the lanes and then shifting down from there instead. Observed reduction in runtime compared to the existing Neon code: Cortex-A510: -41.8% Cortex-A520: -42.6% Cortex-A715: -22.5% Cortex-A720: -22.6% Cortex-X2: -22.7% Cortex-X3: -22.4% Cortex-X4: -19.4% Cortex-X925: -27.0% Bug: b/42280942 Change-Id: I24b092bb352d9858e3d969d82b55940bb00ac7e0 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5802967 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-10-24 21:27:39 +00:00
George Steed	f40042533c	[AArch64] Add SVE2 implementation of I422ToRGB565Row This makes use of the same approach as the Neon code to avoid redundant narrowing and then widening shifts by instead placing the values at the top portion of the lanes and then shifting down from there instead. Observed reduction in runtime compared to the existing Neon code: Cortex-A510: -41.1% Cortex-A520: -38.2% Cortex-A715: -21.5% Cortex-A720: -21.6% Cortex-X2: -21.6% Cortex-X3: -22.0% Cortex-X4: -23.5% Cortex-X925: -21.7% Bug: b/42280942 Change-Id: Id84872141435566bbf94a4bbf0227554b5b5fb91 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5802966 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-10-24 21:27:39 +00:00
George Steed	0dce974ca0	[AArch64] Add SVE2 implementation of I422ToRGB24Row Observed reduction in runtime compared to the existing Neon code: Cortex-A510: -57.8% Cortex-A520: -41.7% Cortex-A715: -28.0% Cortex-A720: -28.1% Cortex-X2: -29.7% Cortex-X3: -28.7% Cortex-X4: -30.5% Cortex-X925: -30.3% Bug: b/42280942 Change-Id: I328bd16babda75fb089c8da8f2714465f658187e Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5802965 Reviewed-by: Frank Barchard <fbarchard@chromium.org> Reviewed-by: Justin Green <greenjustin@google.com>	2024-10-24 02:17:32 +00:00
Frank Barchard	ffd791f749	Check malloc allocation sizes are less than SIZE_MAX Bug: b/371615496 Change-Id: I75a94b08469d6d6b6fd55a8659031cbcb3d48eed Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5912039 Reviewed-by: Wan-Teh Chang <wtc@google.com>	2024-10-07 21:34:15 +00:00
George Steed	dfa279fc65	Re-enable SME when building for AArch64 Android Now that SME has been re-enabled for Linux for a while, also re-enable it for Android when building with a sufficiently new version of LLVM. Bug: b/359006069 Change-Id: Ibaa47e31826cf20136a11d551621fd62c1abab3c Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5908389 Reviewed-by: Frank Barchard <fbarchard@chromium.org> Commit-Queue: Frank Barchard <fbarchard@chromium.org>	2024-10-04 17:43:26 +00:00
George Steed	02c6e8baca	Change ARGBMultiplyRow_C to match Neon The existing behaviour does not round correctly in all cases, so adjust it to match the existing Neon implementation. Update the tests to require bit-exactness and disable other implementations that do not round correctly. Change-Id: Ie790fb4b4805b555d74d689d83802e1dd4f33df5 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5869115 Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-09-23 21:48:33 +00:00
George Steed	a37e6bc81b	[AArch64] Re-enable SME only for Linux and new versions of Clang This was previously disabled in 679e851f653866a49e21f69fe8380bd20123f0ee, so re-enable it but only for Linux where SME is known to work correctly. Change-Id: I2626b03f3854b27162df1b55fc6767e02ffe318d Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5802958 Reviewed-by: Frank Barchard <fbarchard@chromium.org> Reviewed-by: Justin Green <greenjustin@google.com>	2024-09-23 09:29:53 +00:00
George Steed	8315fa1d3a	Avoid duplication of CPU feature disable macros The same conditions are repeated across all *_row.h headers which makes it harder than necessary to guard enabling new architecture features depending on compiler versions etc. Avoid this duplication by merging the conditions into a new cpu_support.h header. Change-Id: Ibe7dfcef138edca6cc36870f1cfbb1bb108083e3 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5802957 Reviewed-by: Frank Barchard <fbarchard@chromium.org> Reviewed-by: Justin Green <greenjustin@google.com>	2024-09-23 09:28:24 +00:00
George Steed	432d186116	[AArch64] Add Neon dot-product implementation for ARGBSepiaRow We can use the dot product instructions to apply the coefficients directly without the need for LD4 de-interleaving load instructions, since these are known to be slow on some micro-architectures. ST4 is also known to be slow on more modern micro-architectures, however avoiding this is left for a future SVE implementation where we can make use of interleaving-narrowing instructions. Reduction in cycle counts observed compared to existing Neon code: Cortex-A55: -5.8% Cortex-A510: -18.9% Cortex-A76: -21.8% Cortex-A720: -30.2% Cortex-X1: -28.6% Cortex-X2: -23.4% Bug: b/42280946 Change-Id: I5887559649cc805a810d867b652c85d48285657d Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5790970 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-09-16 04:31:35 +00:00
George Steed	1c31461771	[AArch64] Add Neon dot-product implementation for ARGBGrayRow We can use dot product instructions to apply the coefficients without needing to use LD4 deinterleaving load instructions, and then TBL to mix in the original alpha component. This is significantly faster on some micro-architectures where LD4 instructions are known to be slow compared to normal loads. Reduction in cycle counts observed compared to existing Neon code: Cortex-A55: -12.6% Cortex-A510: -48.6% Cortex-A76: -39.7% Cortex-A720: -52.3% Cortex-X1: -63.5% Cortex-X2: -67.0% Bug: b/42280946 Change-Id: I3641785e74873438acc00d675f5bc490dfa95b50 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5785972 Reviewed-by: Justin Green <greenjustin@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-09-16 04:31:11 +00:00
Frank Barchard	4620f17058	ScalePlane crash fix for 3/4 scaling - Scaling 48 pixels at a time, but calling code checked for 24 pixels - Added test for scaling to 1080x1920 libyuv_test --gunit_filter=LibYUVScaleTest.I420ScaleTo1080x1920_Box* --libyuv_width=1440 --libyuv_height=2560 Was libyuv_test --gunit_filter=LibYUVScaleTest.I420ScaleTo1080x1920_Box* --libyuv_width=1440 --libyuv_height=2560 [ RUN ] LibYUVScaleTest.I420ScaleTo1080x1920_Box Segmentation fault Traceback (most recent call last): Now [ RUN ] LibYUVScaleTest.I420ScaleTo1080x1920_Box filter 3 - 6741 us C - 3566 us OPT [ OK ] LibYUVScaleTest.I420ScaleTo1080x1920_Box (43 ms) Bug: b/366045177 Change-Id: I0ea6c2d6a32b2e7ca44cd030abc9f248115be44a Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5857554 Reviewed-by: Wan-Teh Chang <wtc@google.com>	2024-09-13 01:20:39 +00:00
Frank Barchard	679e851f65	Convert16To8Row_AVX512BW using vpmovuswb - avx2 is pack/perm is mutating order - cvt method maintains channel order on avx512 Sapphire Rapids Benchmark of 640x360 on Sapphire Rapids AVX512BW [ OK ] LibYUVConvertTest.I010ToNV12_Opt (3547 ms) [ OK ] LibYUVConvertTest.P010ToNV12_Opt (3186 ms) AVX2 [ OK ] LibYUVConvertTest.I010ToNV12_Opt (4000 ms) [ OK ] LibYUVConvertTest.P010ToNV12_Opt (3190 ms) SSE2 [ OK ] LibYUVConvertTest.I010ToNV12_Opt (5433 ms) [ OK ] LibYUVConvertTest.P010ToNV12_Opt (4840 ms) Skylake Xeon Now vpmovuswb [ OK ] LibYUVConvertTest.I010ToNV12_Opt (7946 ms) [ OK ] LibYUVConvertTest.P010ToNV12_Opt (7071 ms) Was vpackuswb [ OK ] LibYUVConvertTest.I010ToNV12_Opt (7684 ms) [ OK ] LibYUVConvertTest.P010ToNV12_Opt (7059 ms) Switch from vpunpcklwd to vpbroadcastw for scale value parameter Was vpunpcklwd %%xmm2,%%xmm2,%%xmm2 vbroadcastss %%xmm2,%%ymm2 Now vpbroadcastw %%xmm2,%%ymm2 Bug: 357439226, 357721018 Change-Id: Ifc9c82ab70dba58af6efa0f57f5f7a344014652e Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5787040 Reviewed-by: Wan-Teh Chang <wtc@google.com>	2024-08-15 20:13:33 +00:00
Wan-Teh Chang	0c2cf03c5c	Fix a -Wundef warning on macOS with Apple silicon Change-Id: Ia78dcc913e06dd8876119a96bd7760c1d2af4341 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5788821 Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-08-14 22:10:43 +00:00
Wan-Teh Chang	02e2ff4745	Note stride params of HalfFloatPlane are in bytes The HalfFloatPlane() function does not follow libyuv's convention of buffer stride in units of the corresponding buffer pointer. Document that. Change-Id: Id8d466ccc2df263a49ad788ab349bc3993a48259 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5770639 Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-08-12 20:17:23 +00:00
Wan-Teh Chang	3cf54e90d3	Fix -Wmissing-prototypes warnings Declare functions as static. Declare functions in a header. Include the header that declares the functions. Delete undeclared and unused functions ScaleFilterRows_NEON() and ScaleRowUp2_16_NEON(). Delete unused function ScaleY() in psnr_main.cc. Change-Id: I182ec30611df83c61ffd01bbab595cd61fb5f1e5 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5778601 Commit-Queue: Wan-Teh Chang <wtc@google.com> Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-08-12 19:08:24 +00:00
Frank Barchard	a97746349b	Add test for I010ToNV12 - Add support for negative height to invert - Fix off by 1 on odd width and height - Bump version to 1895 Initial I010 is 2 step planar conversion libyuv_test '--gunit_filter=*010ToNV12_Opt' --gunit_also_run_disabled_tests --libyuv_width=1280 --libyuv_height=720 --libyuv_repeat=1000 --libyuv_flags=-1 --libyuv_cpu_info=-1 Skylake Xeon [ OK ] LibYUVConvertTest.I010ToNV12_Opt (2675 ms) [ OK ] LibYUVConvertTest.P010ToNV12_Opt (1547 ms) Pixel 7 [ OK ] LibYUVConvertTest.I010ToNV12_Opt (464 ms) [ OK ] LibYUVConvertTest.P010ToNV12_Opt (125 ms) Bug: b/357721018, b/357439226 Change-Id: I2ae59783cf328a6592d0ab80c374ae4dc281daf3 Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5778595 Reviewed-by: Wan-Teh Chang <wtc@google.com>	2024-08-12 18:57:56 +00:00
Chunbo Hua	e23bc72e8e	Bump version number in order to expose new API Bug: 357721018 Change-Id: I2c6e115cd049db2038631195305c5907764d5c7b Reviewed-on: https://chromium-review.googlesource.com/c/libyuv/libyuv/+/5768078 Reviewed-by: Frank Barchard <fbarchard@chromium.org>	2024-08-07 22:10:05 +00:00

1 2 3 4 5 ...

1833 Commits