Experimental
The ElectronDynamicsModels.Experimental submodule holds the batched Tsit5 GPU path (GPUKernelTsit5). It is production staging under active development: retained for electron-batching experiments and not yet validated against the reference solver (the committed accuracy test exercises GPUKernelRK4 only). Prefer the promoted GPU API documented on the Home page.
ElectronDynamicsModels.Experimental.BatchedGPUSplines — Type
BatchedGPUSplines{D, V, M, VO, VK}Flat-storage GPU representation of B trajectory splines. Each electron's data lives at indices offsets[e]+1 : offsets[e]+Ns[e] in the flat arrays. The last row of c1/c2 per electron is unused (interval count = Ns[e]-1) but kept aligned with knot indices for simpler addressing.
ElectronDynamicsModels.Experimental.GPUKernelTsit5 — Type
GPUKernelTsit5Sentinel solver type that selects the batched Tsit5 GPU kernel path.
ElectronDynamicsModels.Experimental.to_batched_cpu — Method
to_batched_cpu(trajs_chunk)Pack B TrajectoryInterpolants into a host-side BatchedGPUSplines (still Vector/Matrix-backed). Adapt.adapt to a backend afterward to upload.
ElectronDynamicsModels.accumulate_potential — Method
accumulate_potential(trajs, screen, ::GPUKernelTsit5, backend; batch_size = 8, n_substeps = 1)Batched GPU path: process electrons in chunks of batch_size, with each kernel launch handling all B electrons in the batch. Compared to the single-electron GPUKernelRK4 path:
- one kernel launch per batch instead of per electron (300/B launches at N_macro = 300, B = 8 ⇒ ~38 launches), amortizing launch + sync overhead;
- flat-with-offsets spline storage handles variable
length(itp.t)per electron with zero padding waste; - 5th-order Tsit5 with FSAL replaces 4th-order RK4, ~2× better effective accuracy per RHS evaluation.
Per-batch device memory is bounded by B × (spline_size + small), so this scales to arbitrary N_macro.
This path is under active development and is not validated against the reference solver. Use GPUKernelRK4 for trusted results.
Arguments
batch_size: B, number of electrons per kernel launch. Pick large enough to saturate GPU occupancy (≈ Nx · Ny · B ≳ 4× SM count × 32 cores) but small enough thatB × spline_sizefits comfortably in VRAM. 8 is a reasonable starting point on a 16 GB card.n_substeps: number of fixed-step Tsit5 sub-steps taken between successive saveat slots (and within the bridge step). Required whenω · δx⁰lies outside Tsit5's accurate range; FSAL is chained across the equal-dtsub-steps so each adds only ~6 RHS evals.