Experimental

The ElectronDynamicsModels.Experimental submodule holds the batched Tsit5 GPU path (GPUKernelTsit5). It is production staging under active development: retained for electron-batching experiments and not yet validated against the reference solver (the committed accuracy test exercises GPUKernelRK4 only). Prefer the promoted GPU API documented on the Home page.

ElectronDynamicsModels.Experimental.BatchedGPUSplines — Type
BatchedGPUSplines{D, V, M, VO, VK}

Flat-storage GPU representation of B trajectory splines. Each electron's data lives at indices offsets[e]+1 : offsets[e]+Ns[e] in the flat arrays. The last row of c1/c2 per electron is unused (interval count = Ns[e]-1) but kept aligned with knot indices for simpler addressing.

source
ElectronDynamicsModels.accumulate_potential — Method
accumulate_potential(trajs, screen, ::GPUKernelTsit5, backend; batch_size = 8, n_substeps = 1)

Batched GPU path: process electrons in chunks of batch_size, with each kernel launch handling all B electrons in the batch. Compared to the single-electron GPUKernelRK4 path:

  • one kernel launch per batch instead of per electron (300/B launches at N_macro = 300, B = 8 ⇒ ~38 launches), amortizing launch + sync overhead;
  • flat-with-offsets spline storage handles variable length(itp.t) per electron with zero padding waste;
  • 5th-order Tsit5 with FSAL replaces 4th-order RK4, ~2× better effective accuracy per RHS evaluation.

Per-batch device memory is bounded by B × (spline_size + small), so this scales to arbitrary N_macro.

Experimental

This path is under active development and is not validated against the reference solver. Use GPUKernelRK4 for trusted results.

Arguments

  • batch_size: B, number of electrons per kernel launch. Pick large enough to saturate GPU occupancy (≈ Nx · Ny · B ≳ 4× SM count × 32 cores) but small enough that B × spline_size fits comfortably in VRAM. 8 is a reasonable starting point on a 16 GB card.
  • n_substeps: number of fixed-step Tsit5 sub-steps taken between successive saveat slots (and within the bridge step). Required when ω · δx⁰ lies outside Tsit5's accurate range; FSAL is chained across the equal-dt sub-steps so each adds only ~6 RHS evals.
source