You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 7518d76
Browse filesBrowse the repository at this point in the historyBrowse files
GPU: pass the work-item indices to the helpers that lack them
MSL has no ambient work-item builtins: the thread and threadgroup positions exist
only as attributes on the kernel entry point, so get_local_id() and its siblings
cannot read them from a device function the way CUDA's threadIdx or OpenCL's
get_local_id() can.
Almost every call site is already inside a function that receives nBlocks,
nThreads, iBlock and iThread, which is how the CPU backend has always worked:
four of the six helpers expand to the bare iBlock and nBlocks there, so they only
compile where those are in scope. Four functions have no index in scope at all
and get them passed in: sortInBlock, GPUTPCCFClusterizer::buildCluster,
GPUTPCCFNoiseSuppression::findMinimaAndPeaks and GPUTPCCFPeakFinder::isPeak.
GPUCA_THREAD_INFO_DECL and GPUCA_THREAD_INFO_PROVIDE gate that so only Metal pays
for it: both expand to nothing on every other backend, and the helpers keep
reading get_local_id(0) rather than open-coding the index arithmetic.
0 commit comments