Commit
This commit does not belong to any branch on this repository, and may belong to a fork outside of the repository.
[TOPI][CUDA] Improve the performance of scatter_nd by:
1. Split into 2 kernels, one does the "Init" and another does the "Update". Thus they can have different Grid/Block configurations to better utilize SMs. 2. Use atomic_add instead of direct assignment, which could avoid the race condtion when multiple indices point to the same location of the output tensor. With this moidification, it's safe now to use more CUDA threads to gain more parallelism.
- Loading branch information