eigen

mirror of https://gitlab.com/libeigen/eigen.git synced 2024-12-21 07:19:46 +08:00

Author	SHA1	Message	Date
Rasmus Munk Larsen	6964ae8d52	Change the sign operator in Eigen to return NaN for NaN arguments, not zero.	2020-07-07 01:54:04 +00:00
David Tellenbach	cb63153183	Make test packetmath C++98 compliant	2020-07-01 20:41:59 +02:00
Sheng Yang	116c5235ac	BF16 for scalar_cmp_with_cast_op	2020-07-01 18:33:42 +00:00
Kan Chen	8731452b97	Delete duplicate test cases in vectorization_logic.cpp	2020-07-01 00:51:15 +00:00
Antonio Sanchez	9cb8771e9c	Fix tensor casts for large packets and casts to/from std::complex The original tensor casts were only defined for `SrcCoeffRatio`:`TgtCoeffRatio` 1:1, 1:2, 2:1, 4:1. Here we add the missing 1:N and 8:1. We also add casting `Eigen::half` to/from `std::complex<T>`, which was missing to make it consistent with `Eigen:bfloat16`, and generalize the overload to work for any complex type. Tests were added to `basicstuff`, `packetmath`, and `cxx11_tensor_casts` to test all cast configurations.	2020-06-30 18:53:55 +00:00
Antonio Sanchez	145e51516f	Fix denormal check pre c++11. `float_denorm_style` is an old-style `enum`, so the `denorm_present` symbol only exists in the `std` namespace prior to c++11.	2020-06-30 17:28:30 +00:00
David Tellenbach	689b57070d	Report custom C++ flags in CMake testing summary	2020-06-30 17:18:54 +00:00
David Tellenbach	f3b8d441f6	Remote CI tags to enable shared runners	2020-06-29 22:15:41 +02:00
Christoph Grüninger	dc0b81fb1d	Pass CMAKE_MAKE_PROGRAM to Fortran language support test Otherwise the Make (or Ninja) program is used, which is installed system wide.	2020-06-27 23:52:38 +02:00
David Tellenbach	13d25f5ed8	Add initial CI configuration file. The initial CI configuration consists of jobs to build and run tests and to build docs.	2020-06-27 00:03:35 +00:00
Antonio Sanchez	7222f0b6b5	Fix packetmath_1 float tests for arm/aarch64. Added missing `pmadd<Packet2f>` for NEON. This leads to significant improvement in precision than previous `pmul+padd`, which was causing the `pcos` tests to fail. Also added an approx test with `std::sin`/`std::cos` since otherwise returning any `a^2+b^2=1` would pass. Modified `log(denorm)` tests. Denorms are not always supported by all systems (returns `::min`), are always flushed to zero on 32-bit arm, and configurably flush to zero on sse/avx/aarch64. This leads to inconsistent results across different systems (i.e. `-inf` vs `nan`). Added a check for existence and exclude ARM. Removed logistic exactness test, since scalar and vectorized versions follow different code-paths due to differences in `pexp` and `pmadd`, which result in slightly different values. For example, exactness always fails on arm, aarch64, and altivec.	2020-06-24 14:03:35 -07:00
Simon Pfreundschuh	14f84978e8	Replaced call to deprecated 'load' function with appropriate call to 'on'.	2020-06-23 11:23:13 +02:00
Antonio Sanchez	ff4e7a0820	Add missing Packet2l/Packet2ul ops for NEON. The current multiply (`pmul`) and comparison operators (`pcmp_lt`, `pcmp_le`, `pcmp_eq`) are missing for packets `Packet2l` and `Packet2ul`. This leads to compile errors for the `packetmath.cpp` tests in clang. Here we add and test the missing ops. Tested: ``` $ aarch64-linux-gnu-g++ -static -I./ '-DEIGEN_TEST_PART_9=1' '-DEIGEN_TEST_PART_10=1' test/packetmath.cpp -o packetmath $ adb push packetmath /data/local/tmp/ $ adb shell "/data/local/tmp/packetmath" $ arm-linux-gnueabihf-g++ -mfpu=neon -static -I./ '-DEIGEN_TEST_PART_9=1' '-DEIGEN_TEST_PART_10=1' test/packetmath.cpp -o packetmath $ adb push packetmath /data/local/tmp/ $ adb shell "/data/local/tmp/packetmath" $ clang++ -target aarch64-linux-android21 -static -I./ '-DEIGEN_TEST_PART_9=1' '-DEIGEN_TEST_PART_10=1' test/packetmath.cpp -o packetmath $ adb push packetmath /data/local/tmp/ $ adb shell "/data/local/tmp/packetmath" $ clang++ -target armv7-linux-android21 -static -mfpu=neon -I./ '-DEIGEN_TEST_PART_9=1' '-DEIGEN_TEST_PART_10=1' test/packetmath.cpp -o packetmath $ adb push packetmath /data/local/tmp/ $ adb shell "/data/local/tmp/packetmath" ```	2020-06-22 11:24:43 -07:00
Antonio Sanchez	03ebdf6acb	Added missing NEON pcasts, update packetmath tests. The NEON `pcast` operators are all implemented and tested for existing packets. This requires adding a `pcast(a,b,c,d,e,f,g,h)` for casting between `int64_t` and `int8_t` in `GenericPacketMath.h`. Removed incorrect `HasHalfPacket` definition for NEON's `Packet2l`/`Packet2ul`. Adjustments were also made to the `packetmath` tests. These include - minor bug fixes for cast tests (i.e. 4:1 casts, only casting for packets that are vectorizable) - added 8:1 cast tests - random number generation - original had uninteresting 0 to 0 casts for many casts between floating-point and integers, and exhibited signed overflow undefined behavior Tested: ``` $ aarch64-linux-gnu-g++ -static -I./ '-DEIGEN_TEST_PART_ALL=1' test/packetmath.cpp -o packetmath $ adb push packetmath /data/local/tmp/ $ adb shell "/data/local/tmp/packetmath" ```	2020-06-21 09:32:31 -07:00
Teng Lu	386d809bde	Support BFloat16 in Eigen	2020-06-20 19:16:24 +00:00
Rasmus Munk Larsen	6b9c92fe7e	Add Apache 2.0 license text in COPYING.APACHE.	2020-06-18 12:45:27 -07:00
Nicolas Mellado	cf7adf3a5d	Update `things you can do` message using cmake commands Print cmake commands instead of make commands, which should work for any generator.	2020-06-16 21:04:33 +00:00
Ilya Tokar	231ce21535	Run two independent chains, when reducing tensors. Running two chains exposes more instruction level parallelism, by allowing to execute both chains at the same time. Results are a bit noisy, but for medium length we almost hit theoretical upper bound of 2x. BM_fullReduction_16T/3 [using 16 threads] 17.3ns ±11% 17.4ns ± 9% ~ (p=0.178 n=18+19) BM_fullReduction_16T/4 [using 16 threads] 17.6ns ±17% 17.0ns ±18% ~ (p=0.835 n=20+19) BM_fullReduction_16T/7 [using 16 threads] 18.9ns ±12% 18.2ns ±10% ~ (p=0.756 n=20+18) BM_fullReduction_16T/8 [using 16 threads] 19.8ns ±13% 19.4ns ±21% ~ (p=0.512 n=20+20) BM_fullReduction_16T/10 [using 16 threads] 23.5ns ±15% 20.8ns ±24% -11.37% (p=0.000 n=20+19) BM_fullReduction_16T/15 [using 16 threads] 35.8ns ±21% 26.9ns ±17% -24.76% (p=0.000 n=20+19) BM_fullReduction_16T/16 [using 16 threads] 38.7ns ±22% 27.7ns ±18% -28.40% (p=0.000 n=20+19) BM_fullReduction_16T/31 [using 16 threads] 146ns ±17% 74ns ±11% -49.05% (p=0.000 n=20+18) BM_fullReduction_16T/32 [using 16 threads] 154ns ±19% 84ns ±30% -45.79% (p=0.000 n=20+19) BM_fullReduction_16T/64 [using 16 threads] 603ns ± 8% 308ns ±12% -48.94% (p=0.000 n=17+17) BM_fullReduction_16T/128 [using 16 threads] 2.44µs ±13% 1.22µs ± 1% -50.29% (p=0.000 n=17+17) BM_fullReduction_16T/256 [using 16 threads] 9.84µs ±14% 5.13µs ±30% -47.82% (p=0.000 n=19+19) BM_fullReduction_16T/512 [using 16 threads] 78.0µs ± 9% 56.1µs ±17% -28.02% (p=0.000 n=18+20) BM_fullReduction_16T/1k [using 16 threads] 325µs ± 5% 263µs ± 4% -19.00% (p=0.000 n=20+16) BM_fullReduction_16T/2k [using 16 threads] 1.09ms ± 3% 0.99ms ± 1% -9.04% (p=0.000 n=20+20) BM_fullReduction_16T/4k [using 16 threads] 7.66ms ± 3% 7.57ms ± 3% -1.24% (p=0.017 n=20+20) BM_fullReduction_16T/10k [using 16 threads] 65.3ms ± 4% 65.0ms ± 3% ~ (p=0.718 n=20+20)	2020-06-16 15:55:11 -04:00
Pedro Caldeira	a475bf14d4	Fix pscatter and pgather for Altivec Complex double	2020-06-16 16:41:02 -03:00
David Tellenbach	c6c84ed961	Fix unused variable warning on Arm	2020-06-15 00:14:58 +02:00
Sebastien Boisvert	6228f27234	Fix #1818 : SparseLU: add methods nnzL() and nnzU() Now this compiles without errors: $ clang++ -I ../../ test_sparseLU.cpp -std=c++03	2020-06-11 23:49:49 +00:00
Sebastien Boisvert	39cbd6578f	Fix #1911 : add benchmark for move semantics with fixed-size matrix $ clang++ -O3 bench/bench_move_semantics.cpp -I. -std=c++11 \ -o bench_move_semantics $ ./bench_move_semantics float copy semantics: 1755.97 ms float move semantics: 55.063 ms double copy semantics: 2457.65 ms double move semantics: 55.034 ms	2020-06-11 23:43:25 +00:00
Antonio Sanchez	a7d2552af8	Remove HasCast and fix packetmath cast tests. The use of the `packet_traits<>::HasCast` field is currently inconsistent with `type_casting_traits<>`, and is unused apart from within `test/packetmath.cpp`. In addition, those packetmath cast tests do not currently reflect how casts are performed in practice: they ignore the `SrcCoeffRatio` and `TgtCoeffRatio` fields, assuming a 1:1 ratio. Here we remove the unsed `HasCast`, and modify the packet cast tests to better reflect their usage.	2020-06-11 17:26:56 +00:00
Sebastien Boisvert	463ec86648	Fix #1757 : remove the word 'suicide'	2020-06-11 00:56:54 +00:00
ShengYang1	b5d66b5e73	Implement scalar_cmp_with_cast_op	2020-06-09 08:12:07 +08:00
Rasmus Munk Larsen	c4059ffcb6	Fix static analyzer warning in SelfadjointProduct.h. Fix compiler warnings in GeneralBlockPanelKernel.h.	2020-06-08 11:48:44 -07:00
Thales Sabino	1fcaaf460f	Update FindComputeCpp.cmake to fix build problems on Windows - Use standard types in SYCL/PacketMath.h to avoid compilation problems on Windows - Add EIGEN_HAS_CONSTEXPR to cxx11_tensor_argmax_sycl.cpp to fix build problems on Windows	2020-06-05 20:51:20 +00:00
David Tellenbach	3ce18d3c8f	Revert ".gitlab-ci.yml: initial commit" This reverts commit `95177362ed` to disable GitLab CI temporarily.	2020-06-05 22:43:49 +02:00
Rasmus Munk Larsen	c2ab36f47a	Fix broken packetmath test for logistic on Arm.	2020-06-04 16:24:47 -07:00
Rasmus Munk Larsen	537e2b322f	Fix typo in previous update to generic predux_any.	2020-06-04 22:25:05 +00:00
Rasmus Munk Larsen	fdc1cbdce3	Avoid implicit float equality comparison in generic predux_any, but use numext::not_equal_strict to avoid breaking builds that compile with -Werror=float-equal.	2020-06-04 22:15:56 +00:00
Rasmus Munk Larsen	daf9bbeca2	Fix compilation error in logistic packet op.	2020-06-03 00:57:41 +00:00
n0mend	6d2a9a524b	Update run instructions for benchCholesky	2020-06-01 18:31:46 +00:00
Gael Guennebaud	029a76e115	Bug #1777 : make the scalar and packet path consistent for the logistic function + respective unit test	2020-05-31 00:53:37 +02:00
Gael Guennebaud	99b7f7cb9c	Fix #556 : warnings with mingw	2020-05-31 00:39:44 +02:00
Gael Guennebaud	72782d13e0	Bug #1767 : increase required cmake version to 3.5.0	2020-05-31 00:31:09 +02:00
Gael Guennebaud	867a756509	Fix #1833 : compilation issue of "array!=scalar" with c++20	2020-05-30 23:53:58 +02:00
Gael Guennebaud	ab615e4114	Save one extra temporary when assigning a sparse product to a row-major sparse matrix	2020-05-30 23:15:12 +02:00
Christoph Junghans	95177362ed	.gitlab-ci.yml: initial commit	2020-05-29 09:23:25 -06:00
Kan Chen	8d1302f566	Add support for PacketBlock<Packet8s,4> and PacketBlock<Packet16uc,4> ptranspose on NEON	2020-05-29 00:33:45 +00:00
Antonio Sánchez	8719b9c5bc	Disable test for 32-bit systems (e.g. ARM, i386) Both i386 and 32-bit ARM do not define __uint128_t. On most systems, if __uint128_t is defined, then so is the macro __SIZEOF_INT128__. https://stackoverflow.com/questions/18531782/how-to-know-if-uint128-t-is-defined1	2020-05-28 17:40:15 +00:00
Yong Tang	8e1df5b082	Fix incorrect usage of `if defined(EIGEN_ARCH_PPC)` => `if EIGEN_ARCH_PPC` This PR tries to fix an incorrect usage of `if defined(EIGEN_ARCH_PPC)` in `Eigen/Core` header. In `Eigen/src/Core/util/Macros.h`, EIGEN_ARCH_PPC was explicitly defined as either 0 or 1. As a result `if defined(EIGEN_ARCH_PPC)` will always be true. This causes issues when building on non PPC platform and `MatrixProduct.h` is not available. This fix changes `if defined(EIGEN_ARCH_PPC)` => `if EIGEN_ARCH_PPC`. Signed-off-by: Yong Tang <yong.tang.github@outlook.com>	2020-05-28 05:53:44 -07:00
Kan Chen	4e7046063b	Fix #1874 : it works on both MSVC 2017 and other platforms.	2020-05-21 18:42:56 +08:00
Pedro Caldeira	2d67af2d2b	Add pscatter for Packet16{u}c (int8)	2020-05-20 17:29:34 -03:00
David Tellenbach	5328cd62b3	Guard usage of decltype since it's a C++11 feature This fixes https://gitlab.com/libeigen/eigen/-/issues/1897	2020-05-20 16:04:16 +02:00
Rasmus Munk Larsen	cc86a31e20	Add guard around specialization for bool, which is only currently implemented for SSE.	2020-05-19 16:21:56 -07:00
Everton Constantino	8a7f360ec3	- Vectorizing MMA packing. - Optimizing MMA kernel. - Adding PacketBlock store to blas_data_mapper.	2020-05-19 19:24:11 +00:00
Rasmus Munk Larsen	a145e4adf5	Add newline at the end of StlIterators.h.	2020-05-15 20:36:00 +00:00
Gael Guennebaud	8ce9630ddb	Fix #1874 : workaround MSVC 2017 compilation issue.	2020-05-15 20:47:32 +02:00
Rasmus Munk Larsen	9b411757ab	Add missing packet ops for bool, and make it pass the same packet op unit tests as other arithmetic types. This change also contains a few minor cleanups: 1. Remove packet op pnot, which is not needed for anything other than pcmp_le_or_nan, which can be done in other ways. 2. Remove the "HasInsert" enum, which is no longer needed since we removed the corresponding packet ops. 3. Add faster pselect op for Packet4i when SSE4.1 is supported. Among other things, this makes the fast transposeInPlace() method available for Matrix<bool>. Run on ************** (72 X 2994 MHz CPUs); 2020-05-09T10:51:02.372347913-07:00 CPU: Intel Skylake Xeon with HyperThreading (36 cores) dL1:32KB dL2:1024KB dL3:24MB Benchmark Time(ns) CPU(ns) Iterations ----------------------------------------------------------------------- BM_TransposeInPlace<float>/4 9.77 9.77 71670320 BM_TransposeInPlace<float>/8 21.9 21.9 31929525 BM_TransposeInPlace<float>/16 66.6 66.6 10000000 BM_TransposeInPlace<float>/32 243 243 2879561 BM_TransposeInPlace<float>/59 844 844 829767 BM_TransposeInPlace<float>/64 933 933 750567 BM_TransposeInPlace<float>/128 3944 3945 177405 BM_TransposeInPlace<float>/256 16853 16853 41457 BM_TransposeInPlace<float>/512 204952 204968 3448 BM_TransposeInPlace<float>/1k 1053889 1053861 664 BM_TransposeInPlace<bool>/4 14.4 14.4 48637301 BM_TransposeInPlace<bool>/8 36.0 36.0 19370222 BM_TransposeInPlace<bool>/16 31.5 31.5 22178902 BM_TransposeInPlace<bool>/32 111 111 6272048 BM_TransposeInPlace<bool>/59 626 626 1000000 BM_TransposeInPlace<bool>/64 428 428 1632689 BM_TransposeInPlace<bool>/128 1677 1677 417377 BM_TransposeInPlace<bool>/256 7126 7126 96264 BM_TransposeInPlace<bool>/512 29021 29024 24165 BM_TransposeInPlace<bool>/1k 116321 116330 6068	2020-05-14 22:39:13 +00:00

1 2 3 4 5 ...

10994 Commits