Motivation
The current MD performance can still be further optimized, especially for GPU inference. Three potential bottlenecks need to be investigated:
- The choice of CUDA
block_size
- The use of
atomicAdd for accumulating energy and virial
- LAMMPS Kokkos interface
Expected outcome
Reduce GPU overhead in MD inference and improve overall simulation speed, especially for large-scale systems where energy and virial accumulation may become performance bottlenecks.
Motivation
The current MD performance can still be further optimized, especially for GPU inference. Three potential bottlenecks need to be investigated:
block_sizeatomicAddfor accumulating energy and virialExpected outcome
Reduce GPU overhead in MD inference and improve overall simulation speed, especially for large-scale systems where energy and virial accumulation may become performance bottlenecks.