xilinx/amd-ross-agentic-ai-assistant
Agent Skills for Developing Products with AMD Embedded Technologies
Convert input code to a multi-stage HLS dataflow, validate with the other hls skills, then hand off to optimize.
Analyze whether a given C/C++ code snippet follows the recommended coding style for array-to-stream conversion in Vitis HLS. Keywords: array, stream, HLS, pragma, ap_fifo, axis, interface
Analyze whether burst inference optimization can be applied to a given C/C++ code snippet in Vitis HLS. Keywords: burst, HLS, pragma, m_axi, AXI, memory, optimization
Get Vitis HLS component basic informations for a given HLS component. The basic informations includes component name, source files, testbench files, include paths and top function
Get Vitis HLS cosimulation report information for a given HLS component.
Analyze whether a given C/C++ code snippet contains a canonical dataflow region for Vitis HLS. Keywords: dataflow, HLS, pragma, canonical, stream, PIPO, hls::task
Get Vitis HLS dataflow process/channel informations for a given HLS component.
Analyze whether a given C/C++ code snippet contains a flattenable loop nest for Vitis HLS.
Diagnose the root cause of an II (initiation interval) violation in a synthesized Vitis HLS component, classify the bottleneck, and cite the supporting evidence — without modifying source or running flows.
Get Vitis HLS implementation report information (post-route clock, resource utilization,Fail Fast table gives utilization percentages) for a given HLS component.
Fill the HLS line-buffer and window-buffer gap for stencil kernels, including TOTAL_ITER drain and verification. Use when generated buffer code misses the drain or the check.
Convert MATLAB sample-based code to plain C++ frame-based loops — analyze algorithm, generate C++ that compiles with g++, verify against MATLAB golden, then hand off to /hls-architect for HLS dataflow architecture.
Iteratively optimize the HLS kernel against a given criteria (e.g. minimize latency, reduce DSP usage, maximize throughput) by applying and measuring pragma and algorithmic changes.
Vitis HLS Performance Pragma — calculate target_ti from a throughput target, cascade through architecture and loops, and present the pragma placement table for user confirmation. Run this before placing any
Decompose a Vitis HLS pipelined region that accesses many streams into a linear chain of smaller sub-pipelines inside a dataflow region, to reduce pipeline-control complexity and improve timing. Keywords: pipeline, dataflow, stream, FIFO, timing, frequency, split pipeline, decompose
Guide users on how to apply pipeline pragma in C/C++ regions, analyze region validity for pipelining, and configure II, rewind, and style options
Defines HLS pragma scope rules. Must be referenced when analyzing HLS designs or generating any code for HLS designs to ensure pragmas are placed in the correct scope.
Run HLS Flow including C Simulation (csim), C Synthesis (csynth), C/RTL Co-Simulation (cosim), Implementation (impl).
Analyze whether stencil pattern optimization can be applied to a given C/C++ code snippet in Vitis HLS. Keywords: stencil, HLS, pragma, array_stencil, window, optimization
Get Vitis HLS synthesis report information (synth report, pragma report, and filtered synthesis log) for a given HLS component.
Analyze whether user code is synthesizable for HLS
Interact with ILA debug cores via vivado-mcp or chipscope-mcp. Discover cores, configure triggers, arm, capture, upload, export waveform data (CSV/VCD), and analyze results. ILA in vivado-mcp supports every device family supported by Vivado. chipscope-mcp supports Versal only. Use when user asks to "capture ILA data", "set ILA trigger", "trigger on signal", "export waveform", "arm ILA", or "trigger immediately".
Interact with VIO debug cores via vivado-mcp or chipscope-mcp. Read input probes, drive output probes, monitor activity, and reset outputs. vivado-mcp supports all device families. chipscope-mcp supports Versal only. Use when user asks to "read VIO inputs", "drive VIO output", "assert <signal> via VIO", "check signal activity", "what is the value of <signal>", or "monitor <probe>".
Guide AI Engine (AIE) integration with Segmented Configuration. Use when designing or building a Versal design that includes AI Engines alongside Segmented Configuration, covering XSA export, PDI packaging, Vitis v++ hardware and emulation flows, and xclbin metadata.
Run Segmented Configuration DRCs and pr_verify design verification checks. Use when validating a design for IO bank conflicts, clocking issues, DDRMC sharing violations, or when comparing routed checkpoints for PL reload compatibility.
Understand Segmented Configuration design constraints for IO banks, CPM/PCIe use cases, and Gen 2 device-specific features (VCU, ISP, ASU, DPDC, 10GbE). Use before finalizing a design that straddles PS and PL domains.
Explain how Segmented Configuration interacts with Dynamic Function eXchange (DFX) on AMD Versal devices. Use this if someone asks about using DFX with Segmented Configuration.
Program Versal hardware using Vivado Hardware Manager in Segmented Configuration mode. Use when programming boot.pdi and pld.pdi through the Vivado GUI, programming configuration flash memory, or enabling the segmented configuration dialog for first-generation Versal devices.
Explain Segmented Configuration concepts for AMD Versal devices — what it is, why it exists, the two-phase boot model, supported device families, and how it compares to standard boot. Use this if someone asks what Segmented Configuration is, how it works, or whether it applies to their device.
Guides the Versal Segmented Configuration PL reload workflow for creating and validating multiple independently compiled pld.pdi variants while keeping boot.pdi fixed. Use when a user asks how to set up a golden PL reload design, spawn a new PL variant project, preserve NoC boot paths with write_noc_solution, size address apertures for reload compatibility, run pr_verify -segcfg_only, or debug unique_id / parent_unique_id mismatches rejected by PLM.
Performance optimization guide for custom AIE kernel development with Peano. Covers restrict pointers, DM bank annotations, loop pragmas, loop hints, software pipelining (pre-RA vs post-RA), loop versioning, function structure, pointer increments, sub-32-bit limitations, vector alignment, type conversion chains, and reading backend optimization hints from the compiler (remarks via -Rpass*/-fsave-optimization-record and warnings such as -Wpass-failed and the aie-multi-slot-pseudo missing-memory-bank hint). Consult when writing or reviewing AIE kernel C++ code for throughput, when diagnosing pipelining failures or high II, or when interpreting build-log warnings and missed-opportunity remarks from the AIE backend.
Comprehensive AIE kernel development guide covering architecture-level optimization techniques, result analysis, and compiler-independent patterns. Covers loop splitting for register pressure, in-place updates, signedness optimization, VLIW slot awareness, result analysis (disassembly, DM maps, calltree, profiling), and general kernel coding practices. Ships aie_api_references.md, the detailed AIE API primitive catalog (per-dtype support, vectorization, accumulator types, math-pattern to API lookup). For Peano compiler-specific pragmas, loop hints, software pipelining details, and optimization remarks, see the vai-aie-compiler-oriented-optimizations skill.
Orchestrate multi-agent parallel development of multiple custom AIE operators for VAIML. Handles workspace partitioning, spawning worker agents, relaying insights between workers, gating integration on verified results, and serializing the final model stitching. Use when developing 2+ custom ops in parallel, when coordinating multi-agent custom op work, or when integrating multiple validated custom ops into a single model.
Find customOperator optimization chains and their ONNX node boundaries
End-to-end custom AIE operator development for VAIML: extract a subgraph from an ONNX model, develop a custom kernel and tiling, validate with x86sim against a CPU reference, and re-integrate the custom op back into the full model. Use when the user wants to create a custom op, write a kernel, define tiling, compile for AIE, simulate with x86sim, or integrate a custom op into a model.
Strip quantization from an ONNX model to produce a non-quantized (float) version suitable for re-quantization. Removes QDQ pairs, Extended QDQ operators, and BFP quantization nodes while preserving the computational graph structure. Supports legacy QDQ (vint8), Extended QDQ (EQDQ/bf16/fp16), and MX6/BFP quantization.
Determine the correct fe_args and vaiml_config flags for a mixed-precision ONNX model based on its quantization pattern. Analyzes model structure to identify island topology, quantization types, and required compiler features. Targets single NPU partition (embedded Telluride scenario).
Identify and quantize Feed-Forward Network (FFN) layers in ONNX models. An FFN is any MatMul+Add (weight+bias) linear layer: MLP projections, attention QKV projections, attention output projections, and classifier heads. This skill detects all FFN patterns, groups them into blocks, and applies per-layer INT8 quantization.
Iteratively improve the VAIML compilation configuration for an ONNX model. Analyzes the model and the flag catalog, proposes a non-default vitisai_config.json, compiles it, validates NPU offload (and board inference time when available), and loops until the user's stop condition is met. Use when the user wants to tune compiler flags, improve NPU offloading, or reduce inference time for a model.
Inspect the result of a VAIML compilation of a mixed-precision ONNX model and reduce the number of NPU partitions toward a single partition. Parses the compiler operation statistics (matched kernels / templated graphs), the partition count, and the AIE/CPU offload split; when more than one partition is produced, reads the CpuBecause messages and compilation warnings, classifies them, and proposes remediation. Remediation prefers the correct frontend (fe_args) flags over model patching; when the partition break originates from quantization, it proposes Quark configuration changes instead. Every proposed change is followed by a re-compilation to evaluate the result.
From L3 DDR DMA-channel profiling, judge compute- vs memory-bound execution and recommend tiling / QoS changes, focused on the channels and operators with the most backpressure/starvation stall
Analyze performance of VAIML compiled models. Uses the AI Analyzer SDK to extract detailed timing, operator metrics, and performance summaries, then answers any question about model performance -- layer timing, custom op cost, bottlenecks, partition breakdown, comparisons between runs, and more. Requires a Ryzen AI / Vitis AI Python environment with dlanalyzer available -- either by sourcing a venv activate script (--t) or by running inside a pre-provisioned environment such as the official Vitis AI Docker container. Use when the user wants to analyze model performance, find slow layers, compare compilation runs, understand custom op overhead, or generate a performance report.
Master orchestrator for mixed-precision ONNX model quantization targeting AMD NPU (Strix/Telluride). Guides users through quantization strategy selection, accuracy validation, compiler flag configuration, and board execution. Supports full bf16, full vint8, hybrid BF16/VINT8 configurations, and edge quantization in runtime. Specialized quantization methods are provided as optional plugin skills discovered at runtime.
Use this skill for configuring ANY Vivado IP block via natural language. Trigger whenever the user asks to configure, customize, parameterize, or set up an IP core — including AXI peripherals, processing systems (Zynq PS, Versal CIPS), memory controllers, DSP blocks, clocking, resets, DMA, interconnects, or any IP with CONFIG.* properties. Also trigger when the user provides a high-level description of desired IP behavior and expects the correct set_property -dict to be generated. This skill uses a two-tier approach: Tier 1 builds the configuration from documentation, Tier 2 recovers from errors using Vivado's own feedback.
Comprehensive Vivado project revision control strategies for standard, DFX (Partial Reconfiguration), IPI Block Design Container, and Segmented Configuration projects. Includes automated tools for project type detection, source analysis, settings capture, and build script generation. Use this skill whenever the user mentions version control, Git, build scripts, project portability, team collaboration, CI/CD for Vivado, exporting sources, preparing a project for handoff, or recreating projects — even if they don't explicitly say "revision control". Also trigger for "make my project portable", "automate project recreation", or "set up Git for my FPGA design".
Analyze Vivado Oasys synthesis elaboration errors, critical warnings, and warnings from synthesis logs and provide actionable RTL code fixes. Log analysis is the primary mode; an active Vivado session is secondary. Inspect the user's actual RTL at every reported file and line before proposing a fix.
Run Vivado's RTL linter to detect design issues before synthesis. Use when user asks to "lint RTL", "check HDL code", "run linter", "find RTL issues", or "check design quality". This skill ONLY analyzes the report generated by the RTL linter command.
Runs and diagnoses FPGA RTL simulations through Vivado using XSim, Questa, ModelSim, VCS, Xcelium, Riviera-PRO, or Active-HDL; supports approval-gated interactive repair; creates self-checking testbenches; runs UVM and seeded regressions; closes XSim code and functional coverage; verifies assertions; and captures SAIF or VCD activity. Use when users ask to simulate HDL, interpret or triage existing simulation logs (including Questa, VCS, Xcelium, or other third-party compile, elaboration, or run logs, and simulator crash or tool-defect reports), debug or repair RTL from waveform evidence, run gate-level or timing simulation, measure coverage, run random regression, analyze power activity, or generate portable stimulus.
Runs 55+ timing methodology checks (UG906) on synthesized FPGA designs to identify timing violations and generate actionable RTL/XDC fixes. Make sure to use this skill whenever user mentions: "timing methodology", "methodology checks", "methodology report", "report_methodology", "UG906", "fix timing issues", "resolve methodology", "timing violations", "methodology violations", "clock violations", or any specific check IDs (TIMING-1 through TIMING-57). Use when user asks "how do I fix", "resolve", "waive", or "understand" any methodology check, warning, or violation. Also use when debugging timing problems after synthesis/implementation, interpreting methodology violation reports, or when user provides violation IDs like "TIMING-1#1". Use even if user just says "I have methodology violations" or "my design has timing warnings".