CUDA — 116 Operations for AI Agents

This page is the canonical reference an AI coding agent uses to refactor, query, and analyze CUDA code through the act MCP server. 116 operations available: 42 refactor, 18 query, 42 analysis, 14 verification. Each operation is callable from Claude Code, Cursor, Codex, OpenCode, or any MCP-compatible agent host. Click any operation for a stable anchor link suitable for citation.

Premium language. Operations on CUDA require an Elite license or above. See pricing.

18Query
42Refactor
42Analysis
14Verify

Worked CUDA examples

act101 applies CUDA-specific correctness and idiom fixes directly to a .cu file. It can add the __device__ qualifier to a plain helper function so kernel code can call it, replace a non-atomic read-modify-write on shared data with the matching atomicAdd call, and insert __syncthreads() between a shared-memory write and the read that consumes it. The atomic-operation conversion addresses a real correctness bug: an unguarded += on data shared across threads is a race condition, where atomicAdd makes the update indivisible. Each example below is the verbatim output of the command shown, run against the file shown.

Add the __device__ qualifier to a helper function

clamp.cu's clampValue is a plain helper that the applyClamp kernel calls directly, so it must be reachable from device code.

$ act refactor-lang add_device_qualifier --file math_helpers.cu --params '{"line":1,"column":1}'

Before

float clampValue(float val, float minVal, float maxVal) {
    if (val < minVal) return minVal;
    if (val > maxVal) return maxVal;
    return val;
}

__global__ void applyClamp(float* data, int n, float lo, float hi) {
    int idx = blockIdx.x * blockDim.x + threadIdx.x;
    if (idx < n) {
        data[idx] = clampValue(data[idx], lo, hi);
    }
}

After

__device__ float clampValue(float val, float minVal, float maxVal) {
    if (val < minVal) return minVal;
    if (val > maxVal) return maxVal;
    return val;
}

__global__ void applyClamp(float* data, int n, float lo, float hi) {
    int idx = blockIdx.x * blockDim.x + threadIdx.x;
    if (idx < n) {
        data[idx] = clampValue(data[idx], lo, hi);
    }
}

clampValue's declaration gains __device__, making it callable from kernel code; the kernel and both function bodies are unchanged.

Convert a shared-memory update to an atomic operation

counter.cu has every thread in countKernel update d_count[0] with a plain +=, which is a data race when multiple threads run concurrently.

$ act refactor-lang convert_to_atomic_operation --file counter.cu --params '{"line":3}'

Before

__global__ void countKernel(int *d_count, int *d_data, int n) {
    int idx = blockIdx.x * blockDim.x + threadIdx.x;
    d_count[0] += d_data[idx];
}

After

__global__ void countKernel(int *d_count, int *d_data, int n) {
    int idx = blockIdx.x * blockDim.x + threadIdx.x;
    atomicAdd(&d_count[0], d_data[idx]);
}

d_count[0] += d_data[idx]; becomes atomicAdd(&d_count[0], d_data[idx]);, so the update is now performed as a single indivisible operation across threads.

Insert __syncthreads() before a shared-memory read

tile.cu's tileKernel writes sdata[threadIdx.x] and then reads it back in the very next line, with no barrier between the block's writers and readers.

$ act refactor-lang add_synchronization --file kernel.cu --params '{"line":4}'

Before

__global__ void tileKernel(float *d_in, float *d_out, int n) {
    __shared__ float sdata[256];
    sdata[threadIdx.x] = d_in[blockIdx.x * blockDim.x + threadIdx.x];
    d_out[blockIdx.x * blockDim.x + threadIdx.x] = sdata[threadIdx.x] * 2.0f;
}

After

__global__ void tileKernel(float *d_in, float *d_out, int n) {
    __shared__ float sdata[256];
    sdata[threadIdx.x] = d_in[blockIdx.x * blockDim.x + threadIdx.x];
    __syncthreads();
    d_out[blockIdx.x * blockDim.x + threadIdx.x] = sdata[threadIdx.x] * 2.0f;
}

__syncthreads(); is inserted between the shared-memory write and the read that consumes it; the rest of the kernel is unchanged.

Query

18 query tools, the same on every supported language. Descriptions live in the shared reference: /docs/query-tools.

callers control_flow data_flow definition diagnostics effect_closure effect_summary fix_auto get_type graph import_organize interface mutations references repo_outline skeleton symbols symbols_batch

Refactor

Operation Description
add-atomic-wrapper Wrap an assignment with a CUDA atomic function
add-bank-conflict-avoidance Insert padding comment to indicate bank conflict avoidance strategy
add-cuda-error-check Insert cudaGetLastError() error check after kernel launch
add-device-qualifier Add __device__ qualifier to function
add-global-qualifier Add __global__ qualifier to function
add-restrict-qualifier Add __restrict__ qualifier to a single pointer parameter
add-restrict-to-parameters Add __restrict__ to all pointer parameters in a kernel signature
add-shared-memory-synchronization Insert __syncthreads() after shared memory writes
add-shared-qualifier Add __shared__ qualifier to variable
add-synchronization Insert __syncthreads() call for thread synchronization
add-warp-convergence-pattern Insert __syncwarp() call for warp-level convergence
convert-global-to-shared Convert a global memory access pattern to use shared memory
convert-to-atomic-operation Convert a simple assignment to an atomic operation
convert-to-constant-memory Convert a variable declaration to use CUDA constant memory
extract-device-function Extract a block of code into a new __device__ function
extract-kernel Extract a block of code into a new __global__ kernel function
extract_function Extract a code selection into a new function — automatically infers parameters, return types, and inserts the call site. Use instead of manually cutting/pasting code. Works without LSP; LSP improves type inference. Params: file (string), new_name (string), start_line (u32), start_column (u32), end_line (u32), end_column (u32) [, preview (bool), receipt (bool)]
extract_variable Extract an expression into a named variable — inserts the declaration and replaces the expression with the variable name. Works without LSP. Params: file (string), new_name (string), start_line (u32), start_column (u32), end_line (u32), end_column (u32) [, preview (bool)]
generate-atomic-operation Generate a CUDA atomic operation call (atomicAdd, atomicMax, etc.)
generate-bank-conflict-avoidance-pattern Generate shared memory declaration with +1 padding to avoid bank conflicts
generate-device-function-stub Generate skeleton __device__ function for GPU execution
generate-error-checking-wrapper Generate cudaGetLastError() error checking wrapper
generate-global-thread-index Generate global thread ID calculation
generate-kernel-launch Generate kernel launch with <<<gridDim, blockDim>>> configuration
generate-kernel-stub Generate skeleton __global__ kernel function
generate-memory-allocation Generate cudaMalloc device memory allocation call
generate-memory-copy-device-to-host Generate cudaMemcpy device-to-host memory copy call
generate-memory-copy-host-to-device Generate cudaMemcpy host-to-device memory copy call
generate-memory-deallocation Generate cudaFree device memory deallocation call
generate-shared-memory-declaration Generate __shared__ memory array declaration
generate-synchronization-barrier Generate __syncthreads() call for thread synchronization
generate-thread-index-variables Generate blockIdx, threadIdx, blockDim variable declarations
inline Inline a variable, function, or method — replace every usage with its definition body, then remove the original. The inverse of extract. Works without LSP (single-file); LSP enables cross-file inlining. Params: file (string), symbol (string) [, line (u32), preview (bool), receipt (bool)]
insert_body Replace a function's implementation body with new code. AST-validated — rejects if the result has parse errors, so you can't accidentally break syntax. Use instead of manual text editing for function rewrites. Params: file (string), symbol (string), code (string) [, commit (bool)]
move_symbol Move a function, class, or type to a different file and automatically update all imports across the codebase. Use instead of manually cut/paste + fixing imports. Works without LSP (single-file); LSP enables cross-file import updates. Params: file (string), symbol (string), destination (string) [, preview (bool), receipt (bool)]
optimize-memory-coalescing Add comment annotation identifying coalescing optimization opportunity
recipe_run Run a codemod recipe: declarative match → transform → optional verify across modeled grammars. Preview lists matches; apply writes with optional E7 receipts and all-or-nothing rollback. Returns a per-site report.
remove-warp-divergence Add comment annotation identifying warp divergence to resolve
rename Rename a symbol and automatically update ALL references across the codebase. Safer and faster than find-and-replace — AST-aware, won't rename strings or comments. Works without LSP (single-file); LSP enables cross-file renames. Params: file (string), old_name (string), new_name (string) [, line (u32), column (u32), preview (bool), receipt (bool)]
rename-device-function Rename a __device__ function and all its call sites
rename-kernel Rename a __global__ kernel function and all its launch sites
simplify-memory-access-pattern Add comment annotation for memory access pattern simplification

Analysis

42 analysis tools, the same on every supported language. Descriptions live in the shared reference: /docs/analysis-tools.

analyze_api_diff analyze_chokepoints analyze_clones analyze_clusters analyze_cohesion analyze_conformance analyze_coupling analyze_cycle_risk analyze_cycles analyze_dead_code analyze_depth analyze_entry_points analyze_export analyze_extraction analyze_fan_balance analyze_features analyze_hotspots analyze_impact analyze_inconsistencies analyze_inheritance analyze_interface_bloat analyze_interfaces analyze_layers analyze_orphan_types analyze_patterns analyze_platform_deps analyze_readiness analyze_roles analyze_seams analyze_stability analyze_surface analyze_test_gaps analyze_thickness analyze_type_completeness churn_hotspots co_change_clusters coverage_overlay ownership_map profile_overlay simulate split_module trace_overlay

Verify

14 verify tools, the same on every supported language. Descriptions live in the shared reference: /docs/verification.

bisect_regression gate generate_test_harness scan secret_surface summarize_pr taint_flow unsafe_surface verify_behavioral_equivalence verify_contract_preserved verify_diff_semantics verify_port_parity verify_side_effects verify_test_impact

← CSSCUE →