eigen

mirror of https://gitlab.com/libeigen/eigen.git synced 2024-12-21 07:19:46 +08:00

Author	SHA1	Message	Date
Eugene Zhulenev	dbca11e880	Remove TensorBlock.h and old TensorBlock/BlockMapper	2019-12-10 14:31:44 -08:00
Deven Desai	c49f0d851a	Fix for HIP breakage detected on 191210 The following commit introduces compile errors when running eigen with hipcc `2918f85ba9` hipcc errors out because it requies the device attribute on the methods within the TensorBlockV2ResourceRequirements struct instroduced by the commit above. The fix is to add the device attribute to those methods	2019-12-10 22:14:05 +00:00
Eugene Zhulenev	2918f85ba9	Do not use std::vector in getResourceRequirements	2019-12-09 16:19:55 -08:00
Artem Belevich	8056a05b54	Undo the block size change. .z is used by the EigenContractionKernelInternal().	2019-12-09 11:10:29 -08:00
Eugene Zhulenev	dbb703d44e	Add async evaluation support to TensorSelectOp	2019-12-09 18:36:13 +00:00
Janek Kozicki	11d6465326	fix AlignedVector3 inconsisent interface with other Vector classes, default constructor and operator- were missing.	2019-12-06 21:07:39 +01:00
Eugene Zhulenev	bb7ccac3af	Add recursive work splitting to EvalShardedByInnerDimContext	2019-12-05 14:51:49 -08:00
Artem Belevich	25230d1862	Improve performance of contraction kernels * Force-inline implementations. They pass around pointers to shared memory blocks. Without inlining compiler must operate via generic pointers. Inlining allows compiler to detect that we're operating on shared memory which allows generation of substantially faster code. * Fixed a long-standing typo which resulted in launching 8x more kernels than we needed (.z dimension of the block is unused by the kernel).	2019-12-05 12:48:34 -08:00
Gael Guennebaud	08eeb648ea	update hg to git hashes	2019-12-05 16:33:24 +01:00
Rasmus Munk Larsen	366cf005b0	Add missing initialization in cxx11_tensor_trace.cpp.	2019-12-04 23:56:37 +00:00
Gael Guennebaud	c488b8b32f	Replace calls to "hg" by calls to "git"	2019-12-04 11:24:06 +01:00
Gael Guennebaud	8fbe0e4699	Update old links to bitbucket to point to gitlab.com	2019-12-04 10:57:07 +01:00
Gael Guennebaud	114a15c66a	Added tag before-git-migration for changeset `a7c7d329d8`	2019-12-04 10:06:00 +01:00
Rasmus Larsen	a7c7d329d8	Merged in ezhulenev/eigen-01 (pull request PR-769) Capture TensorMap by value inside tensor expression AST	2019-12-04 00:49:10 +00:00
Rasmus Larsen	cacf433975	Merged in anshuljl/eigen-2/Anshul-Jaiswal/update-configurevectorizationh-to-not-op-1573079916090 (pull request PR-754) Update ConfigureVectorization.h to not optimize fp16 routines when compiling with cuda. Approved-by: Deven Desai <deven.desai.amd@gmail.com>	2019-12-04 00:45:42 +00:00
Eugene Zhulenev	8f4536e852	Capture TensorMap by value inside tensor expression AST	2019-12-03 16:39:05 -08:00
Rasmus Munk Larsen	4e696901f8	Remove __host__ annotation for device-only function.	2019-12-03 14:33:19 -08:00
Rasmus Munk Larsen	ead81559c8	Use EIGEN_DEVICE_FUNC macro instead of __device__.	2019-12-03 12:08:22 -08:00
Gael Guennebaud	6358599ecb	Fix QuaternionBase::cast for quaternion map and wrapper.	2019-12-03 14:51:14 +01:00
Gael Guennebaud	7745f69013	bug #1776 : fix vector-wise STL iterator's operator-> using a proxy as pointer type. This changeset fixes also the value_type definition.	2019-12-03 14:40:15 +01:00
Rasmus Munk Larsen	66f07efeae	Revert the specialization for scalar_logistic_op<float> introduced in: `77b447c24e` While providing a 50% speedup on Haswell+ processors, the large relative error outside [-18, 18] in this approximation causes problems, e.g., when computing gradients of activation functions like softplus in neural networks.	2019-12-02 17:00:58 -08:00
Rasmus Larsen	3b15373bb3	Merged in ezhulenev/eigen-02 (pull request PR-767) Fix shadow warnings in AlignedBox and SparseBlock	2019-12-02 18:23:11 +00:00
Deven Desai	312c8e77ff	Fix for the HIP build+test errors. Recent changes have introduced the following build error when compiling with HIPCC --------- unsupported/test/../../Eigen/src/Core/GenericPacketMath.h:254:58: error: 'ldexp': no overloaded function has restriction specifiers that are compatible with the ambient context 'pldexp' --------- The fix for the error is to pick the math function(s) from the global namespace (where they are declared as device functions in the HIP header files) when compiling with HIPCC.	2019-12-02 17:41:32 +00:00
Rasmus Larsen	956131d0e6	Merged in codeplaysoftware/eigen/SYCL-Backend (pull request PR-691) SYCL Backend Approved-by: Rasmus Larsen <rmlarsen@google.com>	2019-11-28 16:19:25 +00:00
Mehdi Goli	00f32752f7	[SYCL] Rebasing the SYCL support branch on top of the Einge upstream master branch. * Unifying all loadLocalTile from lhs and rhs to an extract_block function. * Adding get_tensor operation which was missing in TensorContractionMapper. * Adding the -D method missing from cmake for Disable_Skinny Contraction operation. * Wrapping all the indices in TensorScanSycl into Scan parameter struct. * Fixing typo in Device SYCL * Unifying load to private register for tall/skinny no shared * Unifying load to vector tile for tensor-vector/vector-tensor operation * Removing all the LHS/RHS class for extracting data from global * Removing Outputfunction from TensorContractionSkinnyNoshared. * Combining the local memory version of tall/skinny and normal tensor contraction into one kernel. * Combining the no-local memory version of tall/skinny and normal tensor contraction into one kernel. * Combining General Tensor-Vector and VectorTensor contraction into one kernel. * Making double buffering optional for Tensor contraction when local memory is version is used. * Modifying benchmark to accept custom Reduction Sizes * Disabling AVX optimization for SYCL backend on the host to allow SSE optimization to the host * Adding Test for SYCL * Modifying SYCL CMake	2019-11-28 10:08:54 +00:00
Eugene Zhulenev	82a47338df	Fix shadow warnings in AlignedBox and SparseBlock	2019-11-27 16:22:27 -08:00
Rasmus Munk Larsen	ea51a9eace	Add missing EIGEN_DEVICE_FUNC attribute to template specializations for pexp to fix GPU build.	2019-11-27 10:17:09 -08:00
Rasmus Munk Larsen	5a3ebda36b	Fix warning due to missing cast for exponent arguments for std::frexp and std::lexp.	2019-11-26 16:18:29 -08:00
Rasmus Larsen	2df57be856	Merged in realjhol/eigen/fix-warnings (pull request PR-760) Fix warnings	2019-11-26 23:24:23 +00:00
Eugene Zhulenev	5496d0da0b	Add async evaluation support to TensorReverse	2019-11-26 15:02:24 -08:00
Eugene Zhulenev	bc66c88255	Add async evaluation support to TensorPadding/TensorImagePatch/TensorShuffling	2019-11-26 11:41:57 -08:00
Gael Guennebaud	c79b6ffe1f	Add an explicit example for auto and re-evaluation	2019-11-20 17:31:23 +01:00
Hans Johnson	e78ed6e7f3	COMP: Simplify install commands for Eigen Confirm that install directory is identical before and after this simplifying patch. ```bash hg clone <<Eigen>> mkdir eigen-bld cd eigen-bld cmake ../Eigen -DCMAKE_INSTALL_PREFIX:PATH=/tmp/bef make install find /tmp/pre_eigen_modernize >/tmp/bef # Apply this patch cmake ../Eigen -DCMAKE_INSTALL_PREFIX:PATH=/tmp/aft make install find /tmp/post_eigen_modernize \|sed 's/post_e/pre_e/g' >/tmp/aft diff /tmp/bef /tmp/aft ```	2019-11-17 15:14:25 -06:00
Hans Johnson	9d5cdc98c3	COMP: target_compile_definitions requires cmake 2.8.11 Features committed in 2016 have required cmake verison 2.8.11. `sergiu Tue Nov 22 12:25:06 2016 +0100: target_compile_definitions` Set the minimum cmake version to the minimum version that is capable of compiling or installing the code base.	2019-11-17 14:59:32 -06:00
Gael Guennebaud	e5778b87b9	Fix duplicate symbol linking error.	2019-11-20 17:23:19 +01:00
Joel Holdsworth	86eb41f1cb	SparseRef: Fixed alignment warning on ARM GCC	2019-11-07 14:34:06 +00:00
Anshul Jaiswal	c1a67cb5af	Update ConfigureVectorization.h to not optimize fp16 routines when compiling with cuda.	2019-11-06 22:40:38 +00:00
Rasmus Munk Larsen	cc3d0e6a40	Add EIGEN_HAS_INTRINSIC_INT128 macro Add a new EIGEN_HAS_INTRINSIC_INT128 macro, and use this instead of __SIZEOF_INT128__. This fixes related issues with TensorIntDiv.h when building with Clang for Windows, where support for 128-bit integer arithmetic is advertised but broken in practice.	2019-11-06 14:24:33 -08:00
Rasmus Munk Larsen	ee404667e2	Rollback or PR-746 and partial rollback of `668ab3fc47` . std::array is still not supported in CUDA device code on Windows.	2019-11-05 17:17:58 -08:00
Joel Holdsworth	743c925286	test/packetmath: Silence alignment warnings	2019-11-05 19:06:12 +00:00
Rasmus Larsen	0c9745903a	Merged in ezhulenev/eigen-01 (pull request PR-746) Remove internal::smart_copy and replace with std::copy	2019-11-04 20:18:38 +00:00
Hans Johnson	8c8cab1afd	STYLE: Convert CMake-language commands to lower case Ancient CMake versions required upper-case commands. Later command names became case-insensitive. Now the preferred style is lower-case.	2019-10-31 11:36:37 -05:00
Hans Johnson	6fb3e5f176	STYLE: Remove CMake-language block-end command arguments Ancient versions of CMake required else(), endif(), and similar block termination commands to have arguments matching the command starting the block. This is no longer the preferred style.	2019-10-31 11:36:27 -05:00
Rasmus Munk Larsen	f1e8307308	1. Fix a bug in psqrt and make it return 0 for +inf arguments. 2. Simplify handling of special cases by taking advantage of the fact that the builtin vrsqrt approximation handles negative, zero and +inf arguments correctly. This speeds up the SSE and AVX implementations by ~20%. 3. Make the Newton-Raphson formula used for rsqrt more numerically robust: Before: y = y * (1.5 - x/2 * y^2) After: y = y * (1.5 - y * (x/2) * y) Forming y^2 can overflow for very large or very small (denormalized) values of x, while x*y ~= 1. For AVX512, this makes it possible to compute accurate results for denormal inputs down to ~1e-42 in single precision. 4. Add a faster double precision implementation for Knights Landing using the vrsqrt28 instruction and a single Newton-Raphson iteration. Benchmark results: https://bitbucket.org/snippets/rmlarsen/5LBq9o	2019-11-15 17:09:46 -08:00
Gael Guennebaud	2cb2915f90	bug #1744 : fix compilation with MSVC 2017 and AVX512, plog1p/pexpm1 require plog/pexp, but the later was disabled on some compilers	2019-11-15 13:39:51 +01:00
Gael Guennebaud	c3f6fcf2c0	bug #1747 : one more fix for MSVC regarding the Bessel implementation.	2019-11-15 11:12:35 +01:00
Gael Guennebaud	b9837ca9ae	bug #1281 : fix AutoDiffScalar's make_coherent for nested expression of constant ADs.	2019-11-14 14:58:08 +01:00
Gael Guennebaud	0fb6e24408	Fix case issue with Lapack unit tests	2019-11-14 14:16:05 +01:00
Gael Guennebaud	8af045a287	bug #1774 : fix VectorwiseOp::begin()/end() return types regarding constness.	2019-11-14 11:45:52 +01:00
Sakshi Goynar	75b4c0a3e0	PR 751: Fixed compilation issue when compiling using MSVC with /arch:AVX512 flag	2019-10-31 16:09:16 -07:00

1 2 3 4 5 ...

10801 Commits