ffmpeg

Author	SHA1	Message	Date
Michael Niedermayer	8372aaf721	Merge commit '017a06a9ee86b047079166c2694c9c655ff03356' * commit '017a06a9ee86b047079166c2694c9c655ff03356': x86: dsputil: Use correct file name as multiple inclusion guard Merged-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-20 14:58:04 +01:00
Diego Biurrun	017a06a9ee	x86: dsputil: Use correct file name as multiple inclusion guard	2014-02-20 04:16:15 -08:00
Michael Niedermayer	130c33af35	Merge commit 'b23bc95920e2f10b9621857e829c45b064f356c0' * commit 'b23bc95920e2f10b9621857e829c45b064f356c0': x86: dca: Add missing multiple inclusion guards Merged-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-19 15:44:48 +01:00
Diego Biurrun	b23bc95920	x86: dca: Add missing multiple inclusion guards	2014-02-19 10:19:15 +01:00
Hendrik Leppkes	7716eda0aa	vp9/x86: set correct number of registers used in intra pred asm Reviewed-by: "Ronald S. Bultje" <rsbultje@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-18 17:20:14 +01:00
James Almer	07b4b0ca62	tta/x86: add ff_ttafilter_process_dec_{ssse3, sse4} Results are from a Win64 build running on an AMD FX 6300 1121 decicycles in ttafilter_process_dec_c, 16777112 runs, 104 skips 522 decicycles in ff_ttafilter_process_dec_ssse3, 16777149 runs, 67 skips 477 decicycles in ff_ttafilter_process_dec_sse4, 16777156 runs, 60 skips Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-17 13:51:19 +01:00
Ronald S. Bultje	fdb093c4e4	vp9/x86: intra prediction SIMD. Partially based on h264_intrapred. (I hope to eventually merge these two intrapred implementations back together.)	2014-02-17 13:39:00 +01:00
James Almer	ec482e738d	x86/fladsp: add missing check to ff_flacdsp_init_x86() Fixes compilation with flac decoder disabled and encoder enabled Signed-off-by: James Almer <jamrial@gmail.com> Reviewed-by: Paul B Mahol <onemda@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-16 12:06:04 +01:00
Michael Niedermayer	d601106ab1	avcodec/x86/lossless_videodsp: fix w type Fixes fate issues on mingw64 Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-15 06:41:38 +01:00
Peter Ross	b8664c9294	avcodec/vp8dsp: add VP7 idct and loop filter Signed-off-by: Peter Ross <pross@xvid.org> Reviewed-by: "Ronald S. Bultje" <rsbultje@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-15 02:15:35 +01:00
James Almer	e87974bc00	flac/x86: add ff_flac_lpc_32_xop() Tested on an AMD FX 6300 679081 decicycles in ff_flac_lpc_32_xop, 32768 runs 774425 decicycles in ff_flac_lpc_32_sse4, 32768 runs Signed-off-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-13 22:14:59 +01:00
James Darnley	623f380a18	lavc: fix flac encoder and decoder dependencies Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-13 21:00:32 +01:00
Michael Niedermayer	df98b36aa6	Merge commit '5c1c6e82261b856214499b9fef3a08bf3ff6e0ae' * commit '5c1c6e82261b856214499b9fef3a08bf3ff6e0ae': dca: include dcadsp.h in {arm,x86}/dca.h for checkheaders Merged-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-08 17:25:31 +01:00
Michael Niedermayer	dd2b330347	Merge commit '0cffd6fff59f192120dc93aa6c3cb8180f5506e3' * commit '0cffd6fff59f192120dc93aa6c3cb8180f5506e3': x86: use the inline int8x8_fmul_int32 only if inline SSE2 is availbale Merged-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-08 17:06:57 +01:00
Janne Grunau	5c1c6e8226	dca: include dcadsp.h in {arm,x86}/dca.h for checkheaders	2014-02-08 13:38:36 +01:00
Janne Grunau	0cffd6fff5	x86: use the inline int8x8_fmul_int32 only if inline SSE2 is availbale Fixes compilation with MSVC. Also does not rely on on earlier config.h include but include it directly.	2014-02-08 12:10:56 +01:00
Clément Bœsch	669d4f9053	x86/vp9lpf: simplify 2nd transpose in 44/48/88/84. For non-avx optims, this saves 8 movs. before: 1785 decicycles in ff_vp9_loop_filter_h_44_16_ssse3, 524129 runs, 159 skips 3327 decicycles in ff_vp9_loop_filter_h_48_16_ssse3, 262116 runs, 28 skips 2712 decicycles in ff_vp9_loop_filter_h_88_16_ssse3, 4193729 runs, 575 skips 3237 decicycles in ff_vp9_loop_filter_h_84_16_ssse3, 524061 runs, 227 skips after: 1768 decicycles in ff_vp9_loop_filter_h_44_16_ssse3, 524062 runs, 226 skips 3310 decicycles in ff_vp9_loop_filter_h_48_16_ssse3, 262107 runs, 37 skips 2719 decicycles in ff_vp9_loop_filter_h_88_16_ssse3, 4193954 runs, 350 skips 3184 decicycles in ff_vp9_loop_filter_h_84_16_ssse3, 524236 runs, 52 skips	2014-02-08 11:10:23 +01:00
Michael Niedermayer	82ae8a44e6	Merge commit '5b59a9fc6152169599561f04b4f66370edda5c9c' * commit '5b59a9fc6152169599561f04b4f66370edda5c9c': x86: dcadsp: implement int8x8_fmul_int32 Merged-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-08 01:20:33 +01:00
Christophe Gisquet	5b59a9fc61	x86: dcadsp: implement int8x8_fmul_int32 For the callable function (as opposed to the inline one): C SSE SSE2 SSE4 Win32: 47 42 29 26 Win64: 30 33 25 23 The SSE version is neither compiled nor set for ARCH_X86_64, as the inlinable function takes over. Signed-off-by: Janne Grunau <janne-libav@jannau.net>	2014-02-07 22:52:40 +01:00
Loren Merritt	9c978f243a	flac/x86: add ff_flac_lpc_32_sse4() benchmarked on sandybridge x86_64: 1358232 decicycles in flac_lpc_32_c 1244575 decicycles in flac_lpc_32_sse4, James Almer's patch 650045 decicycles in flac_lpc_32_sse4, this patch I haven't tested the edgecases such as odd block lengths odd block length tested-by: James Almer <jamrial@gmail.com> Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-02-06 02:51:19 +01:00
Clément Bœsch	d92a725329	x86/vp9lpf: remove 8 SWAPs in 84/48 transpose.	2014-02-05 07:21:13 +01:00
Clément Bœsch	97dde561de	x86/vp9lpf: remove braindead double pxor.	2014-02-05 07:21:11 +01:00
Clément Bœsch	9a3b05b0a9	x86/vp9lpf: save a few mov in flat8in/hev masks calc.	2014-02-05 07:21:09 +01:00
Clément Bœsch	91d85bb167	x86/vp9lpf: add ff_vp9_loop_filter_[vh]_44_16_{sse2,ssse3,avx}.	2014-02-05 07:21:06 +01:00
Michael Niedermayer	de17ccc774	Merge commit '51daafb02eaf96e0743a37ce95a7f5d02c1fa3c2' * commit '51daafb02eaf96e0743a37ce95a7f5d02c1fa3c2': x86: videodsp: Properly mark sse2 instructions in emulated_edge_mc as such. Conflicts: libavcodec/x86/videodsp_init.c See: `1b3a7e1f42` Merged-by: Michael Niedermayer <michaelni@gmx.at>	2014-01-31 14:30:30 +01:00
Clément Bœsch	c5dd73b890	x86/vp9lpf: add ff_vp9_loop_filter_h_{48,84}_16_{sse2,ssse3,avx}(). 5.40s → 5.30s overall decode time with -threads 1 on ped1080p.webm (i7 920, ssse3)	2014-01-30 19:34:13 +01:00
Ronald S. Bultje	9ee9c679a7	x86: videodsp: Fix a bug in a %if statement where we used '%%' instead of '&&'. Signed-off-by: Janne Grunau <janne-libav@jannau.net>	2014-01-30 15:33:23 +01:00
Ronald S. Bultje	51daafb02e	x86: videodsp: Properly mark sse2 instructions in emulated_edge_mc as such. Should fix crashes or corrupt output on pre-SSE2 CPUs when they were using SSE2-code (e.g. AMD Athlon XP 2400+ or Intel Pentium III) in hfix or hvar single-edge (left/right) extension functions. Signed-off-by: Janne Grunau <janne-libav@jannau.net>	2014-01-30 15:30:01 +01:00
James Almer	644c32ea4b	x86/vp9lpf: add ff_vp9_loop_filter_[vh]_88_16_sse2() Similar gains as the ssse3 version once again Signed-off-by: James Almer <jamrial@gmail.com>	2014-01-28 09:30:55 +01:00
Clément Bœsch	222c46c531	x86/vp9lpf: add ff_vp9_loop_filter_[vh]_88_16_{ssse3,avx}. 9680 decicycles in loop_filter_v_88_16_c, 4193765 runs, 539 skips 9233 decicycles in loop_filter_h_88_16_c, 4193751 runs, 553 skips 1929 decicycles in ff_vp9_loop_filter_v_88_16_ssse3, 4194118 runs, 186 skips 2738 decicycles in ff_vp9_loop_filter_h_88_16_ssse3, 4193861 runs, 443 skips 5.978 → 5.417 overall decode time on ped1080p.webm (-threads 1) Adding SSE2 support should be relatively trivial (just a matter of changing the pshufb [mask_mix] with something else), patch welcome.	2014-01-28 07:36:38 +01:00
Clément Bœsch	822385d775	x86/vp9lpf: add a preload system in FILTER_UPDATE. Allow some macro refactoring in filter14().	2014-01-27 22:39:26 +01:00
Clément Bœsch	315b4775ad	x86/vp9lpf: refactor v/h using common macros for P7 to Q7.	2014-01-27 22:39:26 +01:00
Clément Bœsch	5d144086cc	x86/vp9lpf: faster P7..Q7 accesses. Introduce 2 additional registers for stride3 and mstride3 to allow direct accesses (lea drops). 3931 → 3827 decicycles in ff_vp9_loop_filter_v_16_16_ssse3 Also uses defines to clarify the code.	2014-01-27 22:37:42 +01:00
Clément Bœsch	5f4d04d084	x86/lossless_videodsp: silly one-line cosmetic.	2014-01-25 16:24:50 +01:00
Clément Bœsch	5267e85056	x86/lossless_videodsp: use common macro for add and diff int16 loop.	2014-01-25 14:27:37 +01:00
Clément Bœsch	cddbfd2a95	x86/lossless_videodsp: simplify and explicit aligned/unaligned flags	2014-01-25 11:59:43 +01:00
Ronald S. Bultje	c9e6325ed9	vp9/x86: use explicit register for relative stack references. Before this patch, we explicitly modify rsp, which isn't necessarily universally acceptable, since the space under the stack pointer might be modified in things like signal handlers. Therefore, use an explicit register to hold the stack pointer relative to the bottom of the stack (i.e. rsp). This will also clear out valgrind errors about the use of uninitialized data that started occurring after the idct16x16/ssse3 optimizations were first merged.	2014-01-24 19:25:25 -05:00
Ronald S. Bultje	97474d527f	vp9/x86: iwht4x4 (lossless) mmx.	2014-01-24 19:25:25 -05:00
Ronald S. Bultje	d43efa68bd	vp9/x86: 4x4 iadst SIMD (ssse3) variants. Cycle measurements for intra itxfm_4x4_add on ped1080p.webm: idct_idct: 66 -> 67 cycles (noise measurement) idct_iadst: 199 -> 79 cycles iadst_idct: 165 -> 70 cycles iadst_iadst: 183 -> 82 cycles	2014-01-24 19:25:25 -05:00
Ronald S. Bultje	baf47020cd	vp9/x86: 8x8 iadst SIMD (ssse3/avx) variants. Cycle measurements for intra itxfm_8x8_add on ped1080p.webm: idct_idct: 133 -> 135 cycles (noise measurement) idct_iadst: 900 -> 241 cycles iadst_idct: 864 -> 215 cycles iadst_iadst: 973 -> 310 cycles	2014-01-24 19:25:25 -05:00
Michael Niedermayer	e6d1c66d74	avcodec/x86/lossless_videodsp: disable median optimizations for 16bps They only support upto 15bps Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-01-23 01:51:24 +01:00
Michael Niedermayer	eaacfc7dd1	avcodec/lossless_videodsp: Pass AVCodecContext to init Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-01-23 01:43:00 +01:00
Michael Niedermayer	ef00ef7553	avcodec/x86/lossless_videodsp: port sub_hfyu_median_prediction_int16 to yasm Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-01-22 23:27:27 +01:00
Michael Niedermayer	fad49aae28	avcodec/x86/lossless_videodsp: Port sub_hfyu_median_prediction_mmxext to int16 Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-01-22 22:55:49 +01:00
Michael Niedermayer	fee97f25fa	avcodec/x86/lossless_videodsp: port add_hfyu_median_prediction_mmxext to 16bit Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-01-22 21:11:40 +01:00
Michael Niedermayer	631939bde6	avcodec/x86/lossless_videodsp: add diff_int16_mmx/sse2 Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-01-22 19:41:21 +01:00
Reimar Döffinger	76421982d0	lossless_videodsp.asm: fix compilation. Fixes these errors with nasm: libavcodec/x86/lossless_videodsp.asm:86: error: invalid combination of opcode and operands libavcodec/x86/lossless_videodsp.asm:88: error: invalid combination of opcode and operands I don't know whether movd or movq was meant, but either way maskq vs. maskd must match the mov size. Signed-off-by: Reimar Döffinger <Reimar.Doeffinger@gmx.de>	2014-01-21 19:46:02 +01:00
Michael Niedermayer	83b67ca056	avcodec/x86/lossless_videodsp: Port lorens add_hfyu_left_prediction_ssse3/sse4 to 16bit Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-01-21 02:55:41 +01:00
Michael Niedermayer	63d2be7533	avcodec/x86/lossless_videodsp: use SPLATW in add_int16 Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-01-21 02:33:20 +01:00
Michael Niedermayer	f70d7eb20c	Move add/diff_int16 to lossless_videodsp Signed-off-by: Michael Niedermayer <michaelni@gmx.at>	2014-01-20 21:32:47 +01:00

1 2 3 4 5 ...

1449 Commits