CUDA — 116 Operations for AI Agents
This page is the canonical reference an AI coding agent uses to refactor, query, and analyze CUDA code through the act MCP server. 116 operations available: 42 refactor, 18 query, 42 analysis, 14 verification. Each operation is callable from Claude Code, Cursor, Codex, OpenCode, or any MCP-compatible agent host. Click any operation for a stable anchor link suitable for citation.
Premium language. Operations on CUDA require an Elite license or above. See pricing.
Worked CUDA examples
act101 applies CUDA-specific correctness and idiom fixes directly to a .cu file. It can add the __device__ qualifier to a plain helper function so kernel code can call it, replace a non-atomic read-modify-write on shared data with the matching atomicAdd call, and insert __syncthreads() between a shared-memory write and the read that consumes it. The atomic-operation conversion addresses a real correctness bug: an unguarded += on data shared across threads is a race condition, where atomicAdd makes the update indivisible. Each example below is the verbatim output of the command shown, run against the file shown.
Add the __device__ qualifier to a helper function
clamp.cu's clampValue is a plain helper that the applyClamp kernel calls directly, so it must be reachable from device code.
$ act refactor-lang add_device_qualifier --file math_helpers.cu --params '{"line":1,"column":1}'
Before
float clampValue(float val, float minVal, float maxVal) {
if (val < minVal) return minVal;
if (val > maxVal) return maxVal;
return val;
}
__global__ void applyClamp(float* data, int n, float lo, float hi) {
int idx = blockIdx.x * blockDim.x + threadIdx.x;
if (idx < n) {
data[idx] = clampValue(data[idx], lo, hi);
}
}
After
__device__ float clampValue(float val, float minVal, float maxVal) {
if (val < minVal) return minVal;
if (val > maxVal) return maxVal;
return val;
}
__global__ void applyClamp(float* data, int n, float lo, float hi) {
int idx = blockIdx.x * blockDim.x + threadIdx.x;
if (idx < n) {
data[idx] = clampValue(data[idx], lo, hi);
}
}
clampValue's declaration gains __device__, making it callable from kernel code; the kernel and both function bodies are unchanged.
Convert a shared-memory update to an atomic operation
counter.cu has every thread in countKernel update d_count[0] with a plain +=, which is a data race when multiple threads run concurrently.
$ act refactor-lang convert_to_atomic_operation --file counter.cu --params '{"line":3}'
Before
__global__ void countKernel(int *d_count, int *d_data, int n) {
int idx = blockIdx.x * blockDim.x + threadIdx.x;
d_count[0] += d_data[idx];
}
After
__global__ void countKernel(int *d_count, int *d_data, int n) {
int idx = blockIdx.x * blockDim.x + threadIdx.x;
atomicAdd(&d_count[0], d_data[idx]);
}
d_count[0] += d_data[idx]; becomes atomicAdd(&d_count[0], d_data[idx]);, so the update is now performed as a single indivisible operation across threads.
Insert __syncthreads() before a shared-memory read
tile.cu's tileKernel writes sdata[threadIdx.x] and then reads it back in the very next line, with no barrier between the block's writers and readers.
$ act refactor-lang add_synchronization --file kernel.cu --params '{"line":4}'
Before
__global__ void tileKernel(float *d_in, float *d_out, int n) {
__shared__ float sdata[256];
sdata[threadIdx.x] = d_in[blockIdx.x * blockDim.x + threadIdx.x];
d_out[blockIdx.x * blockDim.x + threadIdx.x] = sdata[threadIdx.x] * 2.0f;
}
After
__global__ void tileKernel(float *d_in, float *d_out, int n) {
__shared__ float sdata[256];
sdata[threadIdx.x] = d_in[blockIdx.x * blockDim.x + threadIdx.x];
__syncthreads();
d_out[blockIdx.x * blockDim.x + threadIdx.x] = sdata[threadIdx.x] * 2.0f;
}
__syncthreads(); is inserted between the shared-memory write and the read that consumes it; the rest of the kernel is unchanged.
Query
18 query tools, the same on every supported language. Descriptions live in the shared reference: /docs/query-tools.
callers control_flow data_flow definition diagnostics effect_closure effect_summary fix_auto get_type graph import_organize interface mutations references repo_outline skeleton symbols symbols_batch
Refactor
| Operation | Description |
|---|---|
add-atomic-wrapper |
Wrap an assignment with a CUDA atomic function |
add-bank-conflict-avoidance |
Insert padding comment to indicate bank conflict avoidance strategy |
add-cuda-error-check |
Insert cudaGetLastError() error check after kernel launch |
add-device-qualifier |
Add __device__ qualifier to function |
add-global-qualifier |
Add __global__ qualifier to function |
add-restrict-qualifier |
Add __restrict__ qualifier to a single pointer parameter |
add-restrict-to-parameters |
Add __restrict__ to all pointer parameters in a kernel signature |
add-shared-memory-synchronization |
Insert __syncthreads() after shared memory writes |
add-shared-qualifier |
Add __shared__ qualifier to variable |
add-synchronization |
Insert __syncthreads() call for thread synchronization |
add-warp-convergence-pattern |
Insert __syncwarp() call for warp-level convergence |
convert-global-to-shared |
Convert a global memory access pattern to use shared memory |
convert-to-atomic-operation |
Convert a simple assignment to an atomic operation |
convert-to-constant-memory |
Convert a variable declaration to use CUDA constant memory |
extract-device-function |
Extract a block of code into a new __device__ function |
extract-kernel |
Extract a block of code into a new __global__ kernel function |
extract_function |
Extract a code selection into a new function — automatically infers parameters, return types, and inserts the call site. Use instead of manually cutting/pasting code. Works without LSP; LSP improves type inference. Params: file (string), new_name (string), start_line (u32), start_column (u32), end_line (u32), end_column (u32) [, preview (bool), receipt (bool)] |
extract_variable |
Extract an expression into a named variable — inserts the declaration and replaces the expression with the variable name. Works without LSP. Params: file (string), new_name (string), start_line (u32), start_column (u32), end_line (u32), end_column (u32) [, preview (bool)] |
generate-atomic-operation |
Generate a CUDA atomic operation call (atomicAdd, atomicMax, etc.) |
generate-bank-conflict-avoidance-pattern |
Generate shared memory declaration with +1 padding to avoid bank conflicts |
generate-device-function-stub |
Generate skeleton __device__ function for GPU execution |
generate-error-checking-wrapper |
Generate cudaGetLastError() error checking wrapper |
generate-global-thread-index |
Generate global thread ID calculation |
generate-kernel-launch |
Generate kernel launch with <<<gridDim, blockDim>>> configuration |
generate-kernel-stub |
Generate skeleton __global__ kernel function |
generate-memory-allocation |
Generate cudaMalloc device memory allocation call |
generate-memory-copy-device-to-host |
Generate cudaMemcpy device-to-host memory copy call |
generate-memory-copy-host-to-device |
Generate cudaMemcpy host-to-device memory copy call |
generate-memory-deallocation |
Generate cudaFree device memory deallocation call |
generate-shared-memory-declaration |
Generate __shared__ memory array declaration |
generate-synchronization-barrier |
Generate __syncthreads() call for thread synchronization |
generate-thread-index-variables |
Generate blockIdx, threadIdx, blockDim variable declarations |
inline |
Inline a variable, function, or method — replace every usage with its definition body, then remove the original. The inverse of extract. Works without LSP (single-file); LSP enables cross-file inlining. Params: file (string), symbol (string) [, line (u32), preview (bool), receipt (bool)] |
insert_body |
Replace a function's implementation body with new code. AST-validated — rejects if the result has parse errors, so you can't accidentally break syntax. Use instead of manual text editing for function rewrites. Params: file (string), symbol (string), code (string) [, commit (bool)] |
move_symbol |
Move a function, class, or type to a different file and automatically update all imports across the codebase. Use instead of manually cut/paste + fixing imports. Works without LSP (single-file); LSP enables cross-file import updates. Params: file (string), symbol (string), destination (string) [, preview (bool), receipt (bool)] |
optimize-memory-coalescing |
Add comment annotation identifying coalescing optimization opportunity |
recipe_run |
Run a codemod recipe: declarative match → transform → optional verify across modeled grammars. Preview lists matches; apply writes with optional E7 receipts and all-or-nothing rollback. Returns a per-site report. |
remove-warp-divergence |
Add comment annotation identifying warp divergence to resolve |
rename |
Rename a symbol and automatically update ALL references across the codebase. Safer and faster than find-and-replace — AST-aware, won't rename strings or comments. Works without LSP (single-file); LSP enables cross-file renames. Params: file (string), old_name (string), new_name (string) [, line (u32), column (u32), preview (bool), receipt (bool)] |
rename-device-function |
Rename a __device__ function and all its call sites |
rename-kernel |
Rename a __global__ kernel function and all its launch sites |
simplify-memory-access-pattern |
Add comment annotation for memory access pattern simplification |
Analysis
42 analysis tools, the same on every supported language. Descriptions live in the shared reference: /docs/analysis-tools.
analyze_api_diff analyze_chokepoints analyze_clones analyze_clusters analyze_cohesion analyze_conformance analyze_coupling analyze_cycle_risk analyze_cycles analyze_dead_code analyze_depth analyze_entry_points analyze_export analyze_extraction analyze_fan_balance analyze_features analyze_hotspots analyze_impact analyze_inconsistencies analyze_inheritance analyze_interface_bloat analyze_interfaces analyze_layers analyze_orphan_types analyze_patterns analyze_platform_deps analyze_readiness analyze_roles analyze_seams analyze_stability analyze_surface analyze_test_gaps analyze_thickness analyze_type_completeness churn_hotspots co_change_clusters coverage_overlay ownership_map profile_overlay simulate split_module trace_overlay
Verify
14 verify tools, the same on every supported language. Descriptions live in the shared reference: /docs/verification.
bisect_regression gate generate_test_harness scan secret_surface summarize_pr taint_flow unsafe_surface verify_behavioral_equivalence verify_contract_preserved verify_diff_semantics verify_port_parity verify_side_effects verify_test_impact