eigen

mirror of https://gitlab.com/libeigen/eigen.git synced 2024-12-21 07:19:46 +08:00

Author	SHA1	Message	Date
Christoph Hertzberg	e6667a7060	Fix stupid shadow-warnings (with old clang versions)	2019-05-07 18:32:19 +02:00
Christoph Hertzberg	e54dc24d62	Restore C++03 compatibility	2019-05-07 18:30:44 +02:00
Rasmus Larsen	ac50afaffa	Merged in ezhulenev/eigen-01 (pull request PR-633) Check if gpu_assert was overridden in TensorGpuHipCudaDefines	2019-04-29 16:29:35 +00:00
Eugene Zhulenev	01d7e6ee9b	Check if gpu_assert was overridden in TensorGpuHipCudaDefines	2019-04-25 11:19:17 -07:00
Eugene Zhulenev	8ead5bb3d8	Fix doxygen warnings to enable statis code analysis	2019-04-24 12:42:28 -07:00
Rasmus Munk Larsen	144ca33321	Remove deprecation annotation from typedef Eigen::Index Index, as it would generate too many build warnings.	2019-04-24 08:50:07 -07:00
Eugene Zhulenev	a7b7f3ca8a	Add missing EIGEN_DEPRECATED annotations to deprecated functions and fix few other doxygen warnings	2019-04-23 17:23:19 -07:00
Anuj Rawat	8c7a6feb8e	Adding lowlevel APIs for optimized RHS packet load in TensorFlow SpatialConvolution Low-level APIs are added in order to optimized packet load in gemm_pack_rhs in TensorFlow SpatialConvolution. The optimization is for scenario when a packet is split across 2 adjacent columns. In this case we read it as two 'partial' packets and then merge these into 1. Currently this only works for Packet16f (AVX512) and Packet8f (AVX2). We plan to add this for other packet types (such as Packet8d) also. This optimization shows significant speedup in SpatialConvolution with certain parameters. Some examples are below. Benchmark parameters are specified as: Batch size, Input dim, Depth, Num of filters, Filter dim Speedup numbers are specified for number of threads 1, 2, 4, 8, 16. AVX512: Parameters \| Speedup (Num of threads: 1, 2, 4, 8, 16) ----------------------------\|------------------------------------------ 128, 24x24, 3, 64, 5x5 \|2.18X, 2.13X, 1.73X, 1.64X, 1.66X 128, 24x24, 1, 64, 8x8 \|2.00X, 1.98X, 1.93X, 1.91X, 1.91X 32, 24x24, 3, 64, 5x5 \|2.26X, 2.14X, 2.17X, 2.22X, 2.33X 128, 24x24, 3, 64, 3x3 \|1.51X, 1.45X, 1.45X, 1.67X, 1.57X 32, 14x14, 24, 64, 5x5 \|1.21X, 1.19X, 1.16X, 1.70X, 1.17X 128, 128x128, 3, 96, 11x11 \|2.17X, 2.18X, 2.19X, 2.20X, 2.18X AVX2: Parameters \| Speedup (Num of threads: 1, 2, 4, 8, 16) ----------------------------\|------------------------------------------ 128, 24x24, 3, 64, 5x5 \| 1.66X, 1.65X, 1.61X, 1.56X, 1.49X 32, 24x24, 3, 64, 5x5 \| 1.71X, 1.63X, 1.77X, 1.58X, 1.68X 128, 24x24, 1, 64, 5x5 \| 1.44X, 1.40X, 1.38X, 1.37X, 1.33X 128, 24x24, 3, 64, 3x3 \| 1.68X, 1.63X, 1.58X, 1.56X, 1.62X 128, 128x128, 3, 96, 11x11 \| 1.36X, 1.36X, 1.37X, 1.37X, 1.37X In the higher level benchmark cifar10, we observe a runtime improvement of around 6% for AVX512 on Intel Skylake server (8 cores). On lower level PackRhs micro-benchmarks specified in TensorFlow tensorflow/core/kernels/eigen_spatial_convolutions_test.cc, we observe the following runtime numbers: AVX512: Parameters \| Runtime without patch (ns) \| Runtime with patch (ns) \| Speedup ---------------------------------------------------------------\|----------------------------\|-------------------------\|--------- BM_RHS_NAME(PackRhs, 128, 24, 24, 3, 64, 5, 5, 1, 1, 256, 56) \| 41350 \| 15073 \| 2.74X BM_RHS_NAME(PackRhs, 32, 64, 64, 32, 64, 5, 5, 1, 1, 256, 56) \| 7277 \| 7341 \| 0.99X BM_RHS_NAME(PackRhs, 32, 64, 64, 32, 64, 5, 5, 2, 2, 256, 56) \| 8675 \| 8681 \| 1.00X BM_RHS_NAME(PackRhs, 32, 64, 64, 30, 64, 5, 5, 1, 1, 256, 56) \| 24155 \| 16079 \| 1.50X BM_RHS_NAME(PackRhs, 32, 64, 64, 30, 64, 5, 5, 2, 2, 256, 56) \| 25052 \| 17152 \| 1.46X BM_RHS_NAME(PackRhs, 32, 256, 256, 4, 16, 8, 8, 1, 1, 256, 56) \| 18269 \| 18345 \| 1.00X BM_RHS_NAME(PackRhs, 32, 256, 256, 4, 16, 8, 8, 2, 4, 256, 56) \| 19468 \| 19872 \| 0.98X BM_RHS_NAME(PackRhs, 32, 64, 64, 4, 16, 3, 3, 1, 1, 36, 432) \| 156060 \| 42432 \| 3.68X BM_RHS_NAME(PackRhs, 32, 64, 64, 4, 16, 3, 3, 2, 2, 36, 432) \| 132701 \| 36944 \| 3.59X AVX2: Parameters \| Runtime without patch (ns) \| Runtime with patch (ns) \| Speedup ---------------------------------------------------------------\|----------------------------\|-------------------------\|--------- BM_RHS_NAME(PackRhs, 128, 24, 24, 3, 64, 5, 5, 1, 1, 256, 56) \| 26233 \| 12393 \| 2.12X BM_RHS_NAME(PackRhs, 32, 64, 64, 32, 64, 5, 5, 1, 1, 256, 56) \| 6091 \| 6062 \| 1.00X BM_RHS_NAME(PackRhs, 32, 64, 64, 32, 64, 5, 5, 2, 2, 256, 56) \| 7427 \| 7408 \| 1.00X BM_RHS_NAME(PackRhs, 32, 64, 64, 30, 64, 5, 5, 1, 1, 256, 56) \| 23453 \| 20826 \| 1.13X BM_RHS_NAME(PackRhs, 32, 64, 64, 30, 64, 5, 5, 2, 2, 256, 56) \| 23167 \| 22091 \| 1.09X BM_RHS_NAME(PackRhs, 32, 256, 256, 4, 16, 8, 8, 1, 1, 256, 56) \| 23422 \| 23682 \| 0.99X BM_RHS_NAME(PackRhs, 32, 256, 256, 4, 16, 8, 8, 2, 4, 256, 56) \| 23165 \| 23663 \| 0.98X BM_RHS_NAME(PackRhs, 32, 64, 64, 4, 16, 3, 3, 1, 1, 36, 432) \| 72689 \| 44969 \| 1.62X BM_RHS_NAME(PackRhs, 32, 64, 64, 4, 16, 3, 3, 2, 2, 36, 432) \| 61732 \| 39779 \| 1.55X All benchmarks on Intel Skylake server with 8 cores.	2019-04-20 06:46:43 +00:00
Rasmus Munk Larsen	039ee52125	Tweak cost model for tensor contraction when parallelizing over the inner dimension. https://bitbucket.org/snippets/rmlarsen/MexxLo	2019-04-12 13:35:10 -07:00
Jonathon Koyle	9a3f06d836	Update TheadPoolDevice example to include ThreadPool creation and passing pointer into constructor.	2019-04-10 10:02:33 -06:00
Deven Desai	66a885b61e	adding EIGEN_DEVICE_FUNC to the recently added TensorContractionKernel constructor. Not having the EIGEN_DEVICE_FUNC attribute on it was leading to compiler errors when compiling Eigen in the ROCm/HIP path	2019-04-08 13:45:08 +00:00
Eugene Zhulenev	629ddebd15	Add missing semicolon	2019-04-02 15:04:26 -07:00
Eugene Zhulenev	4e2f6de1a8	Add support for custom packed Lhs/Rhs blocks in tensor contractions	2019-04-01 11:47:31 -07:00
Deven Desai	2dbea5510f	Merged eigen/eigen into default	2019-03-19 16:52:38 -04:00
David Tellenbach	bd9c2ae3fd	Fix include guard comments	2019-03-15 15:29:17 +01:00
Eugene Zhulenev	001f10e3c9	Fix segfaults with cuda compilation	2019-03-11 09:43:33 -07:00
Eugene Zhulenev	899c16fa2c	Fix a bug in TensorGenerator for 1d tensors	2019-03-11 09:42:01 -07:00
Eugene Zhulenev	0f8bfff23d	Fix a data race in NonBlockingThreadPool	2019-03-11 09:38:44 -07:00
Gael Guennebaud	2df4f00246	Change license from LGPL to MPL2 with agreement from David Harmon.	2019-03-07 18:17:10 +01:00
Rasmus Munk Larsen	3c3f639fe2	Merge.	2019-03-06 11:54:30 -08:00
Rasmus Munk Larsen	f4ec8edea8	Add macro EIGEN_AVOID_THREAD_LOCAL to make it possible to manually disable the use of thread_local.	2019-03-06 11:52:04 -08:00
Rasmus Munk Larsen	41cdc370d0	Fix placement of "#if defined(EIGEN_GPUCC)" guard region. Found with -Wundefined-func-template. Author: tkoeppe@google.com	2019-03-06 11:42:22 -08:00
Rasmus Munk Larsen	cc407c9d4d	Fix placement of "#if defined(EIGEN_GPUCC)" guard region. Found with -Wundefined-func-template. Author: tkoeppe@google.com	2019-03-06 11:40:06 -08:00
Eugene Zhulenev	1bc2a0a57c	Add missing return to NonBlockingThreadPool::LocalSteal	2019-03-06 10:49:49 -08:00
Eugene Zhulenev	4e4dcd9026	Remove redundant steal loop	2019-03-06 10:39:07 -08:00
Eugene Zhulenev	25abaa2e41	Check that inner block dimension is continuous	2019-03-05 17:34:35 -08:00
Eugene Zhulenev	5d9a6686ed	Block evaluation for TensorGeneratorOp	2019-03-05 16:35:21 -08:00
Eugene Zhulenev	a407e022e6	Tune tensor contraction threadpool heuristics	2019-03-05 14:19:59 -08:00
Eugene Zhulenev	56c6373f82	Add an extra check for the RunQueue size estimate	2019-03-05 11:51:26 -08:00
Eugene Zhulenev	b1a8627493	Do not create Tensor<const T> in cxx11_tensor_forced_eval test	2019-03-05 11:19:25 -08:00
Eugene Zhulenev	efb5080d31	Do not initialize invalid fast_strides in TensorGeneratorOp	2019-03-04 16:58:49 -08:00
Eugene Zhulenev	b95941e5c2	Add tiled evaluation for TensorForcedEvalOp	2019-03-04 16:02:22 -08:00
Eugene Zhulenev	694084ecbd	Use fast divisors in TensorGeneratorOp	2019-03-04 11:10:21 -08:00
Rasmus Munk Larsen	cf4a1c81fa	Fix specialization for conjugate on non-complex types in TensorBase.h.	2019-03-01 14:21:09 -08:00
Rasmus Munk Larsen	6560692c67	Improve EventCount used by the non-blocking threadpool. The current algorithm requires threads to commit/cancel waiting in order they called Prewait. Spinning caused by that serialization can consume lots of CPU time on some workloads. Restructure the algorithm to not require that serialization and remove spin waits from Commit/CancelWait. Note: this reduces max number of threads from 2^16 to 2^14 to leave more space for ABA counter (which is now 22 bits). Implementation details are explained in comments.	2019-02-22 13:56:26 -08:00
Gael Guennebaud	9ac1634fdf	Fix conversion warnings	2019-02-19 21:59:53 +01:00
Rasmus Munk Larsen	071629a440	Fix incorrect value of NumDimensions in TensorContraction traits. Reported here: #1671	2019-02-19 10:49:54 -08:00
Rasmus Larsen	efeabee445	Merged in ezhulenev/eigen-01 (pull request PR-590) Do not generate no-op cast() and conjugate() expressions	2019-02-14 21:16:12 +00:00
Eugene Zhulenev	7b837559a7	Fix signed-unsigned return in RuqQueue	2019-02-14 10:40:21 -08:00
Eugene Zhulenev	f0d42d2265	Fix signed-unsigned comparison warning in RunQueue	2019-02-14 10:27:28 -08:00
Eugene Zhulenev	106ba7bb1a	Do not generate no-op cast() and conjugate() expressions	2019-02-14 09:51:51 -08:00
Eugene Zhulenev	8c2f30c790	Speedup Tensor ThreadPool RunQueu::Empty()	2019-02-13 10:20:53 -08:00
Eugene Zhulenev	21eb97d3e0	Add PacketConv implementation for non-vectorizable src expressions	2019-02-08 15:47:25 -08:00
Eugene Zhulenev	1e36166ed1	Optimize TensorConversion evaluator: do not convert same type	2019-02-08 15:13:24 -08:00
Steven Peters	953ca5ba2f	Spline.h: fix spelling "spang" -> "span"	2019-02-08 06:23:24 +00:00
Eugene Zhulenev	59998117bb	Don't do parallel_pack if we can use thread_local memory in tensor contractions	2019-02-07 09:21:25 -08:00
Eugene Zhulenev	8491127082	Do not reduce parallelism too much in contractions with small number of threads	2019-02-04 12:59:33 -08:00
Eugene Zhulenev	eb21bab769	Parallelize tensor contraction only by sharding dimension and use 'thread-local' memory for packing	2019-02-04 10:43:16 -08:00
Gael Guennebaud	d586686924	Workaround lack of support for arbitrary packet-type in Tensor by manually loading half/quarter packets in tensor contraction mapper.	2019-01-30 16:48:01 +01:00
Christoph Hertzberg	a7779a9b42	Hide some annoying unused variable warnings in g++8.1	2019-01-29 16:48:21 +01:00

1 2 3 4 5 ...

2684 Commits