AI Agents Meet HPC: Smarter HPC programming environment on Roihu
Large language models and AI coding agents have quickly become part of many researchers' and developers' daily work. They can explain and write code, troubleshoot errors, and help navigate complex documentation.
Also in high-performance computing (HPC) environments, AI agents are becoming more popular. They are useful to help the user develop HPC applications, but they can also be used as a smart user interface to the system, helping the user to run jobs and manage data on the system. Using an HPC environment is inherently complex, and the agents can make usage both easier and more efficient.
At the same time, using these tools in an HPC environment adds additional concerns around security, permissions, infrastructure usage, and responsible usage. CSC has published a detailed usage policy for how users can use agentic AI tools on Roihu.
To make AI-assisted work easier, while adhering to the policy, a dedicated AI agent environment is now available on Roihu. The environment provides ready-to-use AI agents, HPC-specific tools, and built-in expertise that help users work more effectively.
We are also releasing Docs MCP, a semantic search tool that helps AI tools find relevant user documentation from docs.csc.fi. This MCP server is not limited to the Roihu AI agents, but can also be added to most AI chat interfaces. This allows your tool to have up-to-date information on CSC’s services, enabling it to give you more accurate answers.
At this point, the tools are still experimental and under development, and we are happy to receive feedback on the tools to help us develop our AI tool portfolio.
Roihu AI agents that understand the HPC environment
The Roihu AI environment provides containerized access to several popular coding agents, including OpenCode, Claude Code, and Codex. Users can launch an agent directly from their project directory and begin interacting with their code and files through a conversational interface.
The containerized design serves two important purposes. First, it limits the files an agent can access. Second, it reduces unnecessary load on the underlying storage infrastructure, helping ensure that AI-assisted workflows remain compatible with a shared HPC environment.
Rather than providing a generic chatbot experience, the goal is to offer agents that can assist with practical HPC tasks while operating within clear boundaries and permissions.
Extending agents with HPC-aware tools
A key feature of the environment is support for Model Context Protocol (MCP) servers. MCP provides a standard way for AI agents to access external tools and services. The Roihu environment currently includes dedicated integrations for Slurm and CSC documentation.
The Slurm MCP server allows agents to perform common scheduler-related tasks, such as:
- Inspecting running jobs
- Reviewing job history
- Querying cluster status
- Preparing batch scripts
- Submitting and cancelling jobs with explicit user approval
Because the server exposes only a controlled set of operations and caches results, it helps avoid unnecessary load on scheduler services while still enabling useful automation.
In addition, the Docs MCP gives agents direct access to the CSC User Guide, allowing them to answer questions using platform-specific documentation instead of relying solely on general-purpose model knowledge.
Skills: Teaching agents how to work on Roihu
One of the most interesting aspects of the environment is the use of agent skills.
Skills are lightweight Markdown-based instructions that provide agents with additional context, workflows, and domain knowledge. Rather than loading all information into every conversation, skills are invoked when relevant, making them both efficient and scalable.
The default Roihu environment includes skills focused on:
- Software installation and environment management
- Batch and Slurm script creation
- Job efficiency and optimization
These skills help guide the agent toward HPC best practices and common CSC workflows. For example, when a user asks how to install a package, optimize a job submission script, or troubleshoot a scheduler issue, the relevant skill can provide targeted guidance.
Users can also create and share their own skills, allowing research groups and projects to capture local expertise and make it reusable through AI-assisted workflows.
Responsible use remains essential
AI agents can be powerful productivity tools, but they are not autonomous operators.
Every command executed by an agent runs under the user’s account, and users remain fully responsible for the actions taken. The environment is designed to encourage responsible use through permission controls, restricted access, and clear guidance on data privacy considerations.
As with any AI service, users should understand where prompts and files are sent, what terms apply to the model provider they use, and whether their data is appropriate for processing by external AI services.
From experimentation to everyday productivity
Researchers and HPC users often spend significant time navigating documentation, configuring environments, debugging batch jobs, and solving recurring operational problems. AI agents will not replace HPC expertise, but they can make that expertise easier to access.
By combining modern AI agents with HPC-aware tools, documentation access, and reusable skills, the Roihu AI environment provides a practical way to bring AI assistance closer to everyday scientific computing workflows. Whether you are writing your first Slurm script or maintaining a complex software stack, the new environment offers a starting point for exploring how AI can help you work more effectively on CSC systems.
For detailed information and documentation, please see the documentation on AI tools in CSC docs.
We also welcome any feedback; please send your comments to CSC Service Desk.
Authors: Jussi Heikonen, Sebastian von Alfthan