opencv

Author	SHA1	Message	Date
Alexander Alekhin	89bb028bfc	imgproc(ocl): don't use doubles to process float data	2017-09-07 12:42:20 +03:00
Alexander Alekhin	e3b12bdb59	imgproc(ocl): eliminate div by zero in Canny	2017-08-29 19:29:53 +03:00
Rostislav Vasilikhin	4b75be801e	initial version of Lab2RGB_f tetrahedral interpolation written RGB2Lab_f added, bugs fixed, moved to float several bugs fixed LUT fixed, no switch in tetraInterpolate() temporary code; to be removed and rewritten before refactoring extra interpolations removed, some things to do left added Lab2RGB_b +XYZ version, etc. basic version is done, to be sped up tetra refactored interpolations: LUT for weights, refactor., etc. address arithm optimized initial version of vectorized code added (not compiling now) compilation fixed, now segfaults a lot of fixes, vectorization temp. disabled fixed trilinear shift size, max error dropped from 19 to 10 fixed several bugs (255 vs 256, signed vs unsigned, bIdx) minor changes packed: address arithmetics fixed shorter code experiments with pure integer calculations Lab2RGB max error decreased to 2; need to clean the code ready for vectorization; need cleaning vectorized, to be debugged precision fixed, max error is 2 Lab->XYZ shortened minor fixes Lab2RGB_f version fixed, to be completely rewritten using _b code RGB2Lab_f vectorized minors moved to separate file refactored Lab2RGB to float and int versions minor fix Lab2RGB_f vectorized minor refactoring Lab2RGBint refactored: process methods, vectorize by 4 pix Lab2RGB_f int version is done cleanup extra code code copied to color.cpp fixed blue idx bug optimizations enabled when testing; mulFracConst introduced divConst -> mulFracConst calc min time in perf instead of avg minors process() slightly sped up Lab2RGB_f: disabled int version reinterpret added, minor fixes in names some warnings fixed changes transferred to color.cpp RGB2Lab_f code (and trilinear interpolation code) moved to rgb2lab_faster whitespace shift negative fixed more warnings fixed "constant condition" warnings fixed, little speed up minor changes test_photo decolor fixed changes copied to test_lab.cpp idx bounds checking in LUT init several fixes WIP: softfloat almost integrated test_lab partially rewritten to SoftFloat color.cpp rewritten to SoftFloat test_lab.cpp: accuracy code added several fixes RGB2Lab_b testing fixed splineBuild() rewritten to SoftFloat accuracy control improved rounding fixed Luv <=> RGB: rewritten to SoftFloat OCL cvtColor Lab and Lut rewritten to SoftFloat minor fixes refactored to new SoftFloat interface round() -> cvRound, etc. fixed OCL tests softfloat.cpp: internal functions made static, unused ones removed meaningful constants extra lines removed unused function removed unfinished work it works, need to fix TODOs refactoring; more calls rewritten mulFracConst removed constants made bit exact; minors changes moved to color.cpp fixed 1 bug and 4 warnings OCL: fixed constants pow(x, _1_3f) replaced by cubeRoot(x) fixed compilation on MSVC32 magic constants explained file with internal accuracy&speed tests moved to lab_tetra branch	2017-07-17 00:32:30 +03:00
Rostislav Vasilikhin	aa621d6f3c	magic constants explained	2017-07-06 00:30:53 +03:00
Rostislav Vasilikhin	704c688225	OCL code fixed, fix for NEON added	2017-07-05 22:08:49 +03:00
Maksim Shabunin	ce50df564c	Fixed cvtColor OCL compilation issue (BGRA2mBGRA)	2017-04-05 11:48:29 +03:00
Alexander Alekhin	ba8a6e3533	ocl: don't use vload4 for 3 channel images	2017-03-03 19:36:38 +03:00
mshabunin	8c66531c42	imgproc/CLAHE/ocl: Removed unnecessary __local variable	2017-01-13 16:25:43 +03:00
Li Peng	396921dd23	5x5 gaussian blur optimization Add new 5x5 gaussian blur kernel for CV_8UC1 format, it is 50% ~ 70% faster than current ocl kernel in the perf test. Signed-off-by: Li Peng <peng.li@intel.com>	2016-12-06 09:42:37 +08:00
Li Peng	b69cdb2434	Image pyramids upsampling optimization Add new ocl kernel for image pyramids upsampling, It is 35% faster than current OCL kernel in perf test. Signed-off-by: Li Peng <peng.li@intel.com>	2016-12-02 13:54:58 +08:00
Li Peng	2ca5a7e862	more optimization for warpAffine and warpPerspective Add new OpenCL kernels for bicubic interploation, it is 20% faster than current warp image kernel with bicubic interploation. Signed-off-by: Li Peng <peng.li@intel.com>	2016-11-30 15:43:41 +08:00
Vadim Pisarevsky	c47267ef7f	Merge pull request #7538 from Tetragramm:CLAHEfix	2016-11-29 16:42:05 +00:00
Alexander Alekhin	90b52cd9b8	Merge pull request #7726 from pengli:warp_image	2016-11-29 12:19:18 +00:00
Li Peng	b72d196753	optimization for warpAffine and warpPerspective Add new ocl kernels for warpAffine and warpPerspective, The average performance improvemnt is about 30%. The new ocl kernels require CV_8UC1 format and support nearest neighbor and bilinear interpolation. Signed-off-by: Li Peng <peng.li@intel.com>	2016-11-29 14:55:58 +08:00
Rostislav Vasilikhin	7db43f9fff	fixed wrong equivalence in YUV conversion (#7481 ) * fixed wrong equivalence in YUV conversion * fixed channel order from YVU to YUV	2016-11-23 17:39:18 +03:00
Li Peng	6cb73356b1	laplacian ocl kernel optimization This ocl kernel is 46%~171% faster than current laplacian 3x3 ocl kernel in the perf test, with image format "CV_8UC1". Signed-off-by: Li Peng <peng.li@intel.com>	2016-11-17 12:01:02 +08:00
Li Peng	8d4a7d3dcc	sobel and scharr ocl kernel optimization It improves 108%~230% performance in the perf test with image format "CV_8UC1" and kernel size 3. Signed-off-by: Li Peng <peng.li@intel.com>	2016-11-14 15:34:59 +08:00
Alexander Alekhin	17ffb28807	Merge pull request #7602 from mshabunin:fix-opencl-warnings	2016-11-09 12:35:00 +00:00
Li Peng	8f63f51e81	gaussian blur ocl kernel optimization This ocl kernel is for 3x3 kernel size and CV_8UC1 format It is 115% ~ 300% faster than current ocl path in perf test python ./modules/ts/misc/run.py -t imgproc --gtest_filter=OCL_GaussianBlurFixture* Signed-off-by: Li Peng <peng.li@intel.com>	2016-11-08 11:22:26 +08:00
Alexander Alekhin	442380bfac	Merge pull request #7585 from pengli:morph_filter	2016-11-07 17:11:32 +00:00
mshabunin	3e28d51779	Fixed several OpenCL compiler warnings	2016-11-07 16:49:12 +03:00
Li Peng	35198b84a4	morph ocl kernel for erode and dilate filter This kernel is for CV_8UC1 format and 3x3 kernel size, It is about 33% ~ 55% faster than current ocl kernel with below perf test python ./modules/ts/misc/run.py -t imgproc --gtest_filter=OCL_ErodeFixture* python ./modules/ts/misc/run.py -t imgproc --gtest_filter=OCL_DilateFixture* Also add accuracy test cases for this kernel, the test command is ./bin/opencv_test_imgproc --gtest_filter=OCL_Filter/MorphFilter3x3* Signed-off-by: Li Peng <peng.li@intel.com>	2016-11-04 12:24:24 +08:00
Tetragramm	17df65e666	Fix the OpenCL portion to match the c++ code. Fix an undiscovered bug in the c++ code.	2016-11-03 20:41:16 -05:00
Vadim Pisarevsky	2b7866f21b	Merge pull request #7503 from pengli:box_filter_v2	2016-10-29 21:20:06 +00:00
Li Peng	3607da9f6b	ocl kernel performance optimization for box filter The optimization is for CV_8UC1 format and 3x3 box filter, it is 15%~87% faster than current ocl kernel with below perf test ./modules/ts/misc/run.py -t imgproc --gtest_filter=OCL_BlurFixture* Also add test cases for this ocl kernel. Signed-off-by: Li Peng <peng.li@intel.com>	2016-10-26 11:56:11 +08:00
LukeZhu	ef47bcc88b	Fix the problem: filterSmall.cl report error with double	2016-10-17 15:12:42 +08:00
Alexander Alekhin	b8e08d5d3c	ocl: fix Canny for Intel devices There is an issue with processing of abs(short) function for negative argument. Affected OpenCL devices: - iGPU: Intel(R) HD Graphics 520 (OpenCL 2.0 ) - CPU: Intel(R) Core(TM) i5-6300U CPU @ 2.40GHz (OpenCL 2.0 (Build 10094))	2016-08-09 12:48:06 +03:00
ohnozzy	db9f611767	Add OpenCL support to linearPolar & logPolar Add OpenCL support to linearPolar & logPolar. The OpenCL code use float instead of double, so that it does not require cl_khr_fp64 extension, with slight precision lost. Add explicit conversion Add explicit conversion from double to float to eliminate warning during compilation.	2016-04-24 08:37:56 +08:00
Zhigang Gong	0b08d2559e	fix potential race condition in canny.cl. See the below code snippet: while(l_counter != 0) { int mod = l_counter % LOCAL_TOTAL; int pix_per_thr = l_counter / LOCAL_TOTAL + ((lid < mod) ? 1 : 0); for (int i = 0; i < pix_per_thr; ++i) { int index = atomic_dec(&l_counter) - 1; .... } .... barrier(CLK_LOCAL_MEM_FENCE); } If we don't put a barrier before the for loop, then there is a possiblity that some work item enter this loop but the others are not, the the l_counter will be reduced in the for loop and may be changed to zero, and the other work items may can't enter the while loop. If this happens, it breaks the barrier's rule which requires all the work items reach the same barrier. And it may hang the GPU depends on the implementation of opencl platform. This issue is raised at: https://github.com/Itseez/opencv/issues/5175 Signed-off-by: Zhigang Gong <zhigang.gong@linux.intel.com>	2016-03-15 19:11:15 +08:00
Zhigang Gong	0f7de40e66	Fixed the race condition between inc and dec on the l_counter. Signed-off-by: Zhigang Gong <zhigang.gong@intel.com>	2015-05-26 22:06:18 +08:00
Zhigang Gong	3c85200989	Avoid negative index for a local buffer in Canny.cl. int pix_per_thr = l_counter / LOCAL_TOTAL + ((lid < mod) ? 1 : 0); The pix_per_thr * LOCAL_TOTAL may be larger than l_counter. Thus the index of l_stack may be negative which may cause serious problems. Let's skip the loop when we get negative index and we need to add back the lcounter to keep its balance and avoid potential negative counter. Signed-off-by: Zhigang Gong <zhigang.gong@intel.com>	2015-05-26 08:48:05 +08:00
Pavel Rojtberg	1ea41e7246	fix gftt opencv kernel when using mask	2015-04-22 16:13:50 +02:00
Yan Wang	6e7050555e	Optimize pyrUp_unrolled() by mad function. It could improve performance when image size is large. E.g. OCL_PyrUpFixture_PyrUp.PyrUp/18	2014-11-26 16:55:08 +08:00
Alexander Alekhin	569a95e9e1	Merge pull request #3394 from akarsakov:ocl_canny	2014-11-07 12:42:24 +00:00
Alexander Karsakov	7c870014fb	Correctly unrolled some cycles	2014-11-07 12:13:00 +03:00
Alexander Karsakov	0ec0aeb7d0	Minor optimization for ocl_canny	2014-11-06 13:07:33 +03:00
vbystricky	957e5ef8eb	Fix OpenCL version of HoughLinesP function	2014-11-05 14:31:06 +04:00
Alexander Karsakov	643c906f3d	Added optimized loading to YUV2RGB_422 kernel	2014-10-28 15:07:51 +03:00
Alexander Karsakov	1466621f99	Added loading 4 pixels in line instead of 2 to RGB[A] -> YUV(420) kernel	2014-10-27 16:00:34 +03:00
Alexander Karsakov	60367907fe	Used direct float calculations	2014-10-21 17:18:03 +03:00
Alexander Karsakov	5aa9ac9a77	Added OCL code for YUV422 -> RGB[A]\|BGR[A] color conversion	2014-10-21 17:18:03 +03:00
Alexander Karsakov	c8707b891b	Added OCL code for RGB[A]\|BGR[A] -> YUV_[YV12\|IYUV] color conversion	2014-10-21 17:18:03 +03:00
Alexander Karsakov	1cc17a7186	Added OCL code for YUV2BGR_YV12 and YUV2BGR_IYUV color conversions	2014-10-21 17:18:02 +03:00
Alexander Karsakov	85b60ee3cb	Added support for YUV2RGB[A]_NV21 and YUV2BGR[A]_NV21 conversion	2014-10-21 17:18:02 +03:00
Vadim Pisarevsky	397870d7a5	Merge pull request #3279 from akarsakov:ocl_houghlines	2014-10-09 14:56:45 +00:00
Alexander Karsakov	66a8acfd3d	Optimization for HoughLinesP	2014-10-07 17:53:33 +04:00
Alexander Alekhin	14d5358982	Merge pull request #3210 from akarsakov:ocl_gftt_opt	2014-10-07 09:06:54 +00:00
Alexander Karsakov	eaf5a163b1	Added HoughLinesP OCL implementation	2014-09-29 16:48:16 +04:00
Alexander Karsakov	3695a31606	Combined counter and corner buffers into one	2014-09-29 11:10:57 +04:00
Vadim Pisarevsky	470f427a95	Merge pull request #3232 from Chuanbo-Weng:master	2014-09-18 11:48:29 +00:00

1 2 3 4 5 ...

283 Commits