eigen

mirror of https://gitlab.com/libeigen/eigen.git synced 2025-03-07 18:27:40 +08:00

Author	SHA1	Message	Date
Benoit Steiner	8910442e19	Fixed memcpy, memcpyHostToDevice and memcpyDeviceToHost for Sycl.	2016-12-16 15:45:04 -08:00
Luke Iwanski	54db66c5df	struct -> class in order to silence compilation warning.	2016-12-16 20:25:20 +00:00
Mehdi Goli	35bae513a0	Converting all parallel for lambda to functor in order to prevent kernel duplication name error; adding tensorConcatinationOp backend for sycl.	2016-12-16 19:46:45 +00:00
Mehdi Goli	c5e8546306	Adding asynchandler to sycl queue as lack of it can cause undefined behaviour.	2016-12-15 16:59:57 +00:00
Benoit Steiner	2c2e218471	Avoid using #define since they can conflict with user code	2016-12-14 19:49:15 -08:00
Benoit Steiner	3beb180ee5	Don't call EnvThread::OnCancel by default since it doesn't do anything.	2016-12-14 18:33:39 -08:00
Benoit Steiner	9ff5d0f821	Merged eigen/eigen into default	2016-12-14 17:32:16 -08:00
Mehdi Goli	730eb9fe1c	Adding asynchronous execution as it improves the performance.	2016-12-14 17:38:53 +00:00
Mehdi Goli	2d4a091beb	Adding tensor contraction operation backend for Sycl; adding test for contractionOp sycl backend; adding temporary solution to prevent memory leak in buffer; cleaning up cxx11_tensor_buildins_sycl.h	2016-12-14 15:30:37 +00:00
Benoit Steiner	a432fc102d	Moved the choice of ThreadPool to unsupported/Eigen/CXX11/ThreadPool	2016-12-12 15:24:16 -08:00
Benoit Steiner	8ae68924ed	Made ThreadPoolInterface::Cancel() an optional functionality	2016-12-12 11:58:38 -08:00
Benoit Steiner	76fca22134	Use a more accurate timer to sleep on Linux systems.	2016-12-09 15:12:24 -08:00
Benoit Steiner	4deafd35b7	Introduce a portable EIGEN_SLEEP macro.	2016-12-09 14:52:15 -08:00
Benoit Steiner	aafa97f4d2	Fixed build error with MSVC	2016-12-09 14:42:32 -08:00
Benoit Steiner	2f5b7a199b	Reworked the threadpool cancellation mechanism to not depend on pthread_cancel since it turns out that pthread_cancel doesn't work properly on numerous platforms.	2016-12-09 13:05:14 -08:00
Benoit Steiner	3d59a47720	Added a message to ease the detection of platforms on which thread cancellation isn't supported.	2016-12-08 14:51:46 -08:00
Benoit Steiner	28ee8f42b2	Added a Flush method to the RunQueue	2016-12-08 14:07:56 -08:00
Benoit Steiner	69ef267a77	Added the new threadpool cancel method to the threadpool interface based class.	2016-12-08 14:03:25 -08:00
Benoit Steiner	7bfff85355	Added support for thread cancellation on Linux	2016-12-08 08:12:49 -08:00
Benoit Steiner	462c28e77a	Merged in srvasude/eigen (pull request PR-265) Add Expm1 support to Eigen.	2016-12-05 02:31:11 +00:00
Gael Guennebaud	4465d20403	Add missing generic load methods.	2016-12-03 21:25:04 +01:00
Srinivas Vasudevan	218764ee1f	Added support for expm1 in Eigen.	2016-12-02 14:13:01 -08:00
Mehdi Goli	592acc5bfa	Makingt default numeric_list works with sycl.	2016-12-02 17:58:30 +00:00
Mehdi Goli	79aa2b784e	Adding sycl backend for TensorPadding.h; disbaling __unit128 for sycl in TensorIntDiv.h; disabling cashsize for sycl in tensorDeviceDefault.h; adding sycl backend for StrideSliceOP ; removing sycl compiler warning for creating an array of size 0 in CXX11Meta.h; cleaning up the sycl backend code.	2016-12-01 13:02:27 +00:00
Benoit Steiner	a70393fd02	Cleaned up forward declarations	2016-11-30 21:59:07 -08:00
Benoit Steiner	e073de96dc	Moved the MemCopyFunctor back to TensorSyclDevice since it's the only caller and it makes TensorFlow compile again	2016-11-30 21:36:52 -08:00
Benoit Steiner	fca27350eb	Added the deallocate_all() method back	2016-11-30 20:45:20 -08:00
Benoit Steiner	e633a8371f	Simplified includes	2016-11-30 20:21:18 -08:00
Benoit Steiner	7cd33df4ce	Improved formatting	2016-11-30 20:20:44 -08:00
Benoit Steiner	fd1dc3363e	Merged eigen/eigen into default	2016-11-30 20:16:17 -08:00
Benoit Steiner	f5107010ee	Udated the Sizes class to work on AMD gpus without requiring a separate implementation	2016-11-30 19:57:28 -08:00
Benoit Steiner	e37c2c52d3	Added an implementation of numeric_list that works with sycl	2016-11-30 19:55:15 -08:00
Benoit Steiner	df3da0780d	Updated customIndices2Array to handle various index sizes.	2016-11-30 09:30:12 -08:00
Luke Iwanski	26fff1c5b1	Added EIGEN_STRONG_INLINE to get_sycl_supported_device().	2016-11-30 16:55:22 +00:00
Mehdi Goli	577ce78085	Adding TensorShuffling backend for sycl; adding TensorReshaping backend for sycl; cleaning up the sycl backend.	2016-11-29 15:30:42 +00:00
Benoit Steiner	3011dc94ef	Call internal::array_prod to compute the total size of the tensor.	2016-11-28 09:00:31 -08:00
Benoit Steiner	02080e2b67	Merged eigen/eigen into default	2016-11-27 07:27:30 -08:00
Benoit Steiner	9fd081cddc	Fixed compilation warnings	2016-11-26 20:22:25 -08:00
Benoit Steiner	9f8fbd9434	Merged eigen/eigen into default	2016-11-26 11:28:25 -08:00
Benoit Steiner	67b2c41f30	Avoided unnecessary type conversion	2016-11-26 11:27:29 -08:00
Benoit Steiner	7fe704596a	Added missing array_get method for numeric_list	2016-11-26 11:26:07 -08:00
Mehdi Goli	7318daf887	Fixing LLVM error on TensorMorphingSycl.h on GPU; fixing int64_t crash for tensor_broadcast_sycl on GPU; adding get_sycl_supported_devices() on syclDevice.h.	2016-11-25 16:19:07 +00:00
Benoit Steiner	7ad37606dd	Fixed the documentation of Scalar Tensors	2016-11-24 12:31:43 -08:00
Gael Guennebaud	308961c05e	Fix compilation.	2016-11-23 22:17:52 +01:00
Mehdi Goli	b8cc5635d5	Removing unsupported device from test case; cleaning the tensor device sycl.	2016-11-23 16:30:41 +00:00
Gael Guennebaud	7f6333c32b	Merged in tal500/eigen-eulerangles (pull request PR-237) Euler angles	2016-11-23 15:17:38 +00:00
Gael Guennebaud	f12b368417	Extend polynomial solver unit tests to complexes	2016-11-23 16:05:45 +01:00
Gael Guennebaud	56e5ec07c6	Automatically switch between EigenSolver and ComplexEigenSolver, and fix a few Real versus Scalar issues.	2016-11-23 16:05:10 +01:00
Gael Guennebaud	9246587122	Patch from Oleg Shirokobrod to extend polynomial solver to complexes	2016-11-23 15:42:26 +01:00
Benoit Steiner	f11da1d83b	Made the QueueInterface thread safe	2016-11-20 13:17:08 -08:00
Benoit Steiner	6d781e3e52	Merged eigen/eigen into default	2016-11-20 10:12:54 -08:00
Benoit Steiner	79a07b891b	Fixed a typo	2016-11-20 07:07:41 -08:00
Benoit Steiner	81151bd474	Fixed merge conflicts	2016-11-19 19:12:59 -08:00
Benoit Steiner	9265ca707e	Made it possible to check the state of a sycl device without synchronization	2016-11-19 10:56:24 -08:00
Benoit Steiner	2d1aec15a7	Added missing include	2016-11-19 08:09:54 -08:00
Luke Iwanski	af67335e0e	Added test for cwiseMin, cwiseMax and operator%.	2016-11-19 13:37:27 +00:00
Benoit Steiner	1bdf1b9ce0	Merged in benoitsteiner/opencl (pull request PR-253) OpenCL improvements	2016-11-19 04:44:43 +00:00
Benoit Steiner	a357fe1fb9	Code cleanup	2016-11-18 16:58:09 -08:00
Benoit Steiner	1c6eafb46b	Updated cxx11_tensor_device_sycl to run only on the OpenCL devices available on the host	2016-11-18 16:43:27 -08:00
Benoit Steiner	ca754caa23	Only runs the cxx11_tensor_reduction_sycl on devices that are available.	2016-11-18 16:31:14 -08:00
Benoit Steiner	dc601d79d1	Added the ability to run test exclusively OpenCL devices that are listed by sycl::device::get_devices().	2016-11-18 16:26:50 -08:00
Benoit Steiner	110b7f8d9f	Deleted unnecessary semicolons	2016-11-18 14:06:17 -08:00
Benoit Steiner	b5e3285e16	Test broadcasting on OpenCL devices with 64 bit indexing	2016-11-18 13:44:20 -08:00
Benoit Steiner	37c2c516a6	Cleaned up the sycl device code	2016-11-18 12:38:06 -08:00
Benoit Steiner	7335c49204	Fixed the cxx11_tensor_device_sycl test	2016-11-18 12:37:13 -08:00
Mehdi Goli	15e226d7d3	adding Benoit changes on the TensorDeviceSycl.h	2016-11-18 16:34:54 +00:00
Mehdi Goli	622805a0c5	Modifying TensorDeviceSycl.h to always create buffer of type uint8_t and convert them to the actual type at the execution on the device; adding the queue interface class to separate the lifespan of sycl queue and buffers,created for that queue, from Eigen::SyclDevice; modifying sycl tests to support the evaluation of the results for both row major and column major data layout on all different devices that are supported by Sycl{CPU; GPU; and Host}.	2016-11-18 16:20:42 +00:00
Luke Iwanski	5159675c33	Added isnan, isfinite and isinf for SYCL device. Plus test for that.	2016-11-18 16:01:48 +00:00
Tal Hadad	76b2a3e6e7	Allow to construct EulerAngles from 3D vector directly. Using assignment template struct to distinguish between 3D vector and 3D rotation matrix.	2016-11-18 15:01:06 +02:00
Luke Iwanski	927bd62d2a	Now testing out (+=, =) in.FUNC() and out (+=, =) out.FUNC()	2016-11-18 11:16:42 +00:00
Benoit Steiner	7c30078b9f	Merged eigen/eigen into default	2016-11-17 22:53:37 -08:00
Benoit Steiner	553f50b246	Added a way to detect errors generated by the opencl device from the host	2016-11-17 21:51:48 -08:00
Benoit Steiner	72a45d32e9	Cleanup	2016-11-17 21:29:15 -08:00
Benoit Steiner	4349fc640e	Created a test to check that the sycl runtime can successfully report errors (like ivision by 0). Small cleanup	2016-11-17 20:27:54 -08:00
Benoit Steiner	a6a3fd0703	Made TensorDeviceCuda.h compile on windows	2016-11-17 16:15:27 -08:00
Benoit Steiner	004344cf54	Avoid calling log(0) or 1/0	2016-11-17 11:56:44 -08:00
Luke Iwanski	7878756dea	Fixed existing test.	2016-11-17 17:46:55 +00:00
Luke Iwanski	c5130dedbe	Specialised basic math functions for SYCL device.	2016-11-17 11:47:13 +00:00
Benoit Steiner	b5c75351e3	Merged eigen/eigen into default	2016-11-14 15:54:44 -08:00
Rasmus Munk Larsen	32df1b1046	Reduce dispatch overhead in parallelFor by only calling thread_pool.Schedule() for one of the two recursive calls in handleRange. This avoids going through the scedule path to push both recursive calls onto another thread-queue in the binary tree, but instead executes one of them on the main thread. At the leaf level this will still activate a full complement of threads, but will save up to 50% of the overhead in Schedule (random number generation, insertion in queue which includes signaling via atomics).	2016-11-14 14:18:16 -08:00
Mehdi Goli	05e8c2a1d9	Adding extra test for non-fixed size to broadcast; Replacing stcl with sycl.	2016-11-14 18:13:53 +00:00
Mehdi Goli	f8ca893976	Adding TensorFixsize; adding sycl device memcpy; adding insial stage of slicing.	2016-11-14 17:51:57 +00:00
Mehdi Goli	a5c3f15682	Adding comment to TensorDeviceSycl.h and cleaning the code.	2016-11-11 19:06:34 +00:00
Mehdi Goli	3be3963021	Adding EIGEN_STRONG_INLINE back; using size() instead of dimensions.TotalSize() on Tensor.	2016-11-10 19:16:31 +00:00
Mehdi Goli	12387abad5	adding the missing in eigen_assert!	2016-11-10 18:58:08 +00:00
Mehdi Goli	2e704d4257	Adding Memset; optimising MecopyDeviceToHost by removing double copying;	2016-11-10 18:45:12 +00:00
Benoit Steiner	75c080b176	Added a test to validate memory transfers between host and sycl device	2016-11-09 06:23:42 -08:00
Benoit Steiner	db3903498d	Merged in benoitsteiner/opencl (pull request PR-246) Improved support for OpenCL	2016-11-08 22:28:44 +00:00
Benoit Steiner	dcc14bee64	Fixed the formatting of the code	2016-11-08 14:24:46 -08:00
Luke Iwanski	912cb3d660	#if EIGEN_EXCEPTION -> #ifdef EIGEN_EXCEPTIONS.	2016-11-08 22:01:14 +00:00
Luke Iwanski	1b345b0895	Fix for SYCL queue initialisation.	2016-11-08 21:56:31 +00:00
Luke Iwanski	1b95717358	Use try/catch only when exceptions are enabled.	2016-11-08 21:08:53 +00:00
Mehdi Goli	d57430dd73	Converting all sycl buffers to uninitialised device only buffers; adding memcpyHostToDevice and memcpyDeviceToHost on syclDevice; modifying all examples to obey the new rules; moving sycl queue creating to the device based on Benoit suggestion; removing the sycl specefic condition for returning m_result in TensorReduction.h according to Benoit suggestion.	2016-11-08 17:08:02 +00:00
Benoit Steiner	ad086b03e4	Removed unnecessary statement	2016-11-05 12:43:27 -07:00
Benoit Steiner	dad177be01	Added missing includes	2016-11-05 10:04:42 -07:00
Gael Guennebaud	55b4fd1d40	Extend mpreal unit test to check LLT with complexes.	2016-11-05 11:28:53 +01:00
Benoit Steiner	d46a36cc84	Merged eigen/eigen into default	2016-11-04 18:22:55 -07:00
Mehdi Goli	0ebe3808ca	Removed the sycl include from Eigen/Core and moved it to Unsupported/Eigen/CXX11/Tensor; added TensorReduction for sycl (full reduction and partial reduction); added TensorReduction test case for sycl (full reduction and partial reduction); fixed the tile size on TensorSyclRun.h based on the device max work group size;	2016-11-04 18:18:19 +00:00
Benoit Steiner	3e37166d0b	Merged in benoitsteiner/opencl (pull request PR-244) Disable vectorization on device only when compiling for sycl	2016-11-02 22:01:03 +00:00
Benoit Steiner	0585b2965d	Disable vectorization on device only when compiling for sycl	2016-11-02 11:44:27 -07:00
Benoit Steiner	e6e77ed08b	Don't call lgamma_r when compiling for an Apple device, since the function isn't available on MacOS	2016-11-02 09:55:39 -07:00
Benoit Steiner	b238f387b4	Pulled latest updates from trunk	2016-11-02 08:53:13 -07:00
Benoit Steiner	c8db17301e	Special functions require math.h: make sure it is included.	2016-11-02 08:51:52 -07:00
Benoit Steiner	e44519744e	Merged in benoitsteiner/opencl (pull request PR-243) Fixed the ambiguity in callig make_tuple for sycl backend.	2016-11-02 02:56:58 +00:00
Rasmus Munk Larsen	0a6ae41555	Merged eigen/eigen into default	2016-11-01 15:37:00 -07:00
Rasmus Munk Larsen	b730952414	Don't attempts to use lgamma_r for CUDA devices. Fix type in lgamma_impl<double>.	2016-11-01 15:34:19 -07:00
Mehdi Goli	51af6ae971	Fixed the ambiguity in callig make_tuple for sycl backend.	2016-10-31 16:35:51 +00:00
Benoit Steiner	0a9ad6fc72	Worked around Visual Studio compilation errors	2016-10-28 07:54:27 -07:00
Benoit Steiner	d5f88e2357	Sharded the tensor_image_patch test to help it run on low power devices	2016-10-27 21:48:21 -07:00
Benoit Steiner	0b4b0f11e8	Fixed a few more compilation warnings	2016-10-28 04:01:01 +00:00
Benoit Steiner	306daa24a3	Fixed a compilation warning	2016-10-28 03:50:31 +00:00
Benoit Steiner	8471cf1996	Fixed compilation warning	2016-10-28 03:46:08 +00:00
Benoit Steiner	b0c5bfdf78	Added missing template parameters	2016-10-28 03:43:41 +00:00
Rasmus Munk Larsen	2ebb314fa7	Use threadsafe versions of lgamma and lgammaf if possible.	2016-10-27 16:17:12 -07:00
Gael Guennebaud	530f20c21a	Workaround MSVC issue.	2016-10-27 21:51:37 +02:00
Benoit Steiner	0a4c4d40b4	Removed a template parameter for fixed sized tensors	2016-10-26 18:47:37 -07:00
Benoit Steiner	5f2dd503ff	Replaced tabs with spaces	2016-10-25 20:40:58 -07:00
Benoit Steiner	1644bafe29	Code cleanup	2016-10-25 20:36:14 -07:00
Benoit Steiner	cf20b30d65	Merge latest updates from trunk	2016-10-20 09:42:05 -07:00
Luke Iwanski	03b63e182c	Added SYCL include in Tensor.	2016-10-20 15:32:44 +01:00
Benoit Steiner	d3943cd50c	Fixed a few typos in the ternary tensor expressions types	2016-10-19 12:56:12 -07:00
Tal Hadad	15eca2432a	Euler tests: Tighter precision when no roll exists and clean code.	2016-10-18 23:24:57 +03:00
Tal Hadad	6f4f12d1ed	Add isApprox() and cast() functions. test cases included	2016-10-17 22:23:47 +03:00
Tal Hadad	7402cfd4cc	Add safty for near pole cases and test them better.	2016-10-17 20:42:08 +03:00
Tal Hadad	58f5d7d058	Fix calc bug, docs and better testing. Test code changes: * better coded * rand and manual numbers * singularity checking	2016-10-16 14:39:26 +03:00
Mehdi Goli	e36cb91c99	Fixing the code indentation in the TensorReduction.h file.	2016-10-14 18:03:00 +01:00
Tal Hadad	078a202621	Merge Hongkai Dai correct range calculation, and remove ranges from API. Docs updated.	2016-10-14 16:03:28 +03:00
Luke Iwanski	e742da8b28	Merged ComputeCpp into default.	2016-10-14 13:36:51 +01:00
Mehdi Goli	524fa4c46f	Reducing the code by generalising sycl backend functions/structs.	2016-10-14 12:09:55 +01:00
Hongkai Dai	014d9f1d9b	implement euler angles with the right ranges	2016-10-13 14:45:51 -07:00
Benoit Steiner	d0ee2267d6	Relaxed the resizing checks so that they don't fail with gcc >= 5.3	2016-10-13 10:59:46 -07:00
Benoit Steiner	7e4a6754b2	Merged eigen/eigen into default	2016-10-12 22:42:33 -07:00
Gael Guennebaud	091d373ee9	Fix outer-stride.	2016-10-12 21:47:52 +02:00
Benoit Steiner	7f0599b6eb	Manually define int16_t and uint16_t when compiling with Visual Studio	2016-10-08 22:56:32 -07:00
Benoit Steiner	5266ff8966	Cleaned up a regression test	2016-10-08 19:12:44 +00:00
Benoit Steiner	5c68051cd7	Merge the content of the ComputeCpp branch into the default branch	2016-10-07 11:04:16 -07:00
RJ Ryan	bfc264abe8	Add a test that GPU complex product reductions match CPU reductions.	2016-10-06 11:10:14 -07:00
RJ Ryan	e2e9cdd169	Fully support complex types in SumReducer and MeanReducer when building for CUDA by using scalar_sum_op and scalar_product_op instead of operator+ and operator*.	2016-10-06 10:49:48 -07:00
Benoit Steiner	d7f9679a34	Fixed a couple of compilation warnings	2016-10-05 15:00:32 -07:00
Benoit Steiner	ae1385c7e4	Pull the latest updates from trunk	2016-10-05 14:54:36 -07:00
Benoit Steiner	73b0012945	Fixed compilation warnings	2016-10-05 14:24:24 -07:00
Benoit Steiner	c84084c0c0	Fixed compilation warning	2016-10-05 14:15:41 -07:00
Benoit Steiner	4387433acf	Increased the robustness of the reduction tests on fp16	2016-10-05 10:42:41 -07:00
Benoit Steiner	aad20d700d	Increase the tolerance to numerical noise.	2016-10-05 10:39:24 -07:00
Benoit Steiner	8b69d5d730	::rand() returns a signed integer on win32	2016-10-05 08:55:02 -07:00
Benoit Steiner	ed7a220b04	Fixed a typo that impacts windows builds	2016-10-05 08:51:31 -07:00
Benoit Steiner	ceee1c008b	Silenced compilation warning	2016-10-04 18:47:53 -07:00
Benoit Steiner	6af5ac7e27	Cleanup the cuda executor code.	2016-10-04 08:52:13 -07:00
Benoit Steiner	2f6d1607c8	Cleaned up the random number generation code.	2016-10-04 08:38:23 -07:00
Benoit Steiner	616a7a1912	Improved support for compiling CUDA code with clang as the host compiler	2016-10-03 17:09:33 -07:00
Benoit Steiner	422530946f	Renamed the SYCL tests to follow the standard naming convention.	2016-09-30 08:22:10 -07:00
Benoit Steiner	2bda1b0d93	Updated the tensor sum and mean reducer to enable them to process complex numbers on cuda gpus.	2016-09-28 17:08:41 -07:00
Mehdi Goli	dd602e62c8	Converting alias template to nested struct in order to be compatible with CXX-03	2016-09-27 16:21:19 +01:00
Benoit Steiner	6565f8d60f	Made the initialization of a CUDA device thread safe.	2016-09-26 11:00:32 -07:00
Benoit Steiner	f6ac51a054	Made TensorEvalTo compatible with c++0x again.	2016-09-23 16:45:17 -07:00
Benoit Steiner	00d4e65f00	Deleted unused TensorMap data member	2016-09-23 16:44:45 -07:00
Benoit Steiner	1301d744f8	Made the gaussian generator usable on GPU	2016-09-22 19:04:44 -07:00
RJ Ryan	608b1acd6d	Don't use c++11 features and fix include.	2016-09-20 07:49:05 -07:00
RJ Ryan	b2c6dc48d9	Add CUDA-specific std::complex<T> specializations for scalar_sum_op, scalar_difference_op, scalar_product_op, and scalar_quotient_op.	2016-09-20 07:18:20 -07:00
Gael Guennebaud	3ada6e4bed	Merged hongkai-dai/eigen/tip into default (bug #1298 )	2016-09-19 22:08:06 +02:00
Benoit Steiner	c3ca9b1e76	Deleted some unecessary and confusing EIGEN_DEVICE_FUNC	2016-09-19 11:33:39 -07:00
Hongkai Dai	5dcc6d301a	remove ternary operator in euler angles	2016-09-19 10:30:30 -07:00
Luke Iwanski	c771df6bc3	Updated the owners of the file.	2016-09-19 14:09:25 +01:00
Luke Iwanski	b91e021172	Merged with default.	2016-09-19 14:03:54 +01:00
Luke Iwanski	cb81975714	Partial OpenCL support via SYCL compatible with ComputeCpp CE.	2016-09-19 12:44:13 +01:00
Emil Fresk	6edd2e2851	Made AutoDiffJacobian more intuitive to use and updated for C++11 Changes: * Removed unnecessary types from the Functor by inferring from its types * Removed inputs() function reference, replaced with .rows() * Updated the forward constructor to use variadic templates * Added optional parameters to the Fuctor for passing parameters, control signals, etc * Has been tested with fixed size and dynamic matricies Ammendment by chtz: overload operator() for compatibility with not fully conforming compilers	2016-09-16 14:03:55 +02:00
Gael Guennebaud	18f6e47815	Fix order of "static inline".	2016-09-16 11:32:54 +02:00
Benoit Steiner	488ad7dd1b	Added missing EIGEN_DEVICE_FUNC qualifiers	2016-09-14 13:35:00 -07:00
Benoit Steiner	e4d4d15588	Register the cxx11_tensor_device only for recent cuda architectures (i.e. >= 3.0) since the test instantiate contractions that require a modern gpu.	2016-09-12 19:01:52 -07:00
Benoit Steiner	4dfd888c92	CUDA contractions require arch >= 3.0: don't compile the cuda contraction tests on older architectures.	2016-09-12 18:49:01 -07:00
Benoit Steiner	028e299577	Fixed a bug impacting some outer reductions on GPU	2016-09-12 18:36:52 -07:00
Benoit Steiner	5f50f12d2c	Added the ability to compute the absolute value of a complex number on GPU, as well as a test to catch the problem.	2016-09-12 13:46:13 -07:00
Benoit Steiner	8321dcce76	Merged latest updates from trunk	2016-09-12 10:33:05 -07:00
Benoit Steiner	eb6ba00cc8	Properly size the list of waiters	2016-09-12 10:31:55 -07:00
Benoit Steiner	a618094b62	Added a resize method to MaxSizeVector	2016-09-12 10:30:53 -07:00
Gael Guennebaud	471eac5399	bug #1195 : move NumTraits::Div<>::Cost to internal::scalar_div_cost (with some specializations in arch/SSE and arch/AVX)	2016-09-08 08:36:27 +02:00
Gael Guennebaud	e1642f485c	bug #1288 : fix memory leak in arpack wrapper.	2016-09-05 18:01:30 +02:00
Gael Guennebaud	dabc81751f	Fix compilation when cuda_fp16.h does not exist.	2016-09-05 17:14:20 +02:00
Benoit Steiner	87a8a1975e	Fixed a regression test	2016-09-02 19:29:33 -07:00
Benoit Steiner	13df3441ae	Use MaxSizeVector instead of std::vector: xcode sometimes assumes that std::vector allocates aligned memory and therefore issues aligned instruction to initialize it. This can result in random crashes when compiling with AVX instructions enabled.	2016-09-02 19:25:47 -07:00
Benoit Steiner	cadd124d73	Pulled latest update from trunk	2016-09-02 15:30:02 -07:00
Benoit Steiner	05b0518077	Made the index type an explicit template parameter to help some compilers compile the code.	2016-09-02 15:29:34 -07:00
Benoit Steiner	adf864fec0	Merged in rmlarsen/eigen (pull request PR-222) Fix CUDA build broken by changes to min and max reduction.	2016-09-02 14:11:20 -07:00
Rasmus Munk Larsen	13e93ca8b7	Fix CUDA build broken by changes to min and max reduction.	2016-09-02 13:41:36 -07:00
Benoit Steiner	6c05c3dd49	Fix the cxx11_tensor_cuda.cu test on 32bit platforms.	2016-09-02 11:12:16 -07:00
Benoit Steiner	039e225f7f	Added a test for nullary expressions on CUDA Also check that we can mix 64 and 32 bit indices in the same compilation unit	2016-09-01 13:28:12 -07:00
Benoit Steiner	c53f783705	Updated the contraction code to support constant inputs.	2016-09-01 11:41:27 -07:00
Gael Guennebaud	46475eff9a	Adjust Tensor module wrt recent change in nullary functor	2016-09-01 13:40:45 +02:00
Gael Guennebaud	72a4d49315	Fix compilation with CUDA 8	2016-09-01 13:39:33 +02:00
Rasmus Munk Larsen	a1e092d1e8	Fix bugs to make min- and max reducers with correctly with IEEE infinities.	2016-08-31 15:04:16 -07:00
Gael Guennebaud	1f84f0d33a	merge EulerAngles module	2016-08-30 10:01:53 +02:00
Gael Guennebaud	e074f720c7	Include missing forward declaration of SparseMatrix	2016-08-29 18:56:46 +02:00
Gael Guennebaud	6cd7b9ea6b	Fix compilation with cuda 8	2016-08-29 11:06:08 +02:00
Gael Guennebaud	35a8e94577	bug #1167 : simplify installation of header files using cmake's install(DIRECTORY ...) command.	2016-08-29 10:59:37 +02:00
Gael Guennebaud	0f56b5a6de	enable vectorization path when testing half on cuda, and add test for log1p	2016-08-26 14:55:51 +02:00
Gael Guennebaud	965e595f02	Add missing log1p method	2016-08-26 14:55:00 +02:00
Benoit Steiner	34ae80179a	Use array_prod instead of calling TotalSize since TotalSize is only available on DSize.	2016-08-15 10:29:14 -07:00
Benoit Steiner	fe73648c98	Fixed a bug in the documentation.	2016-08-12 10:00:43 -07:00
Benoit Steiner	e3a8dfb02f	std::erfcf doesn't exist: use numext::erfc instead	2016-08-11 15:24:06 -07:00
Benoit Steiner	64e68cbe87	Don't attempt to optimize partial reductions when the optimized implementation doesn't buy anything.	2016-08-08 19:29:59 -07:00

... 2 3 4 5 6 ...

2329 Commits