Agentic systems for science, and the cyberinfrastructure they run on.
AI that helps run the facility and AI that carries a physicist's analysis, built on facility context and earned trust.
One credential-brokered Model Context Protocol endpoint through which any AI assistant on a physicist's laptop reaches the Analysis Facility. The broker authenticates every caller against the facility's Keycloak, mints short-lived per-user credentials, applies policy and audits every call, so raw grid credentials never reach the assistant. Built by Giordon Stark, it fronts Rucio, AMI, HTCondor, JupyterLab and the AF filesystem for about 800 users, with a self-service portal, and any analysis facility can deploy it.
Project page →Elwood orchestrates AI agents that carry out physics analysis workflows on the Analysis Facility, from Rucio data discovery through ServiceX delivery and Coffea processing to pyhf statistical inference. It is designed for facility reach and operates under Shannon, the Lab's governance framework for AI-assisted operations, where AI proposes and humans decide. Elwood is part of the IRIS-HEP agentic-analysis roadmap and the CLARIPHY community effort on AI agentic analysis systems.
AEGIS is an agentic operations platform that provides operational intelligence for ATLAS Distributed Computing Operations. Agents reason over monitoring data, job and transfer analytics and site status to detect problems, explain them and propose actions to operators, improving the monitoring and operation of the distributed computing systems used by ATLAS and potentially other LHC experiments.
AEGIS on GitHub →Shannon is the Lab's framework for letting agents help run a facility without being handed the keys. Principles, scoped policies and runbooks live in a git repository and are reviewed as pull requests; a runbook is the atomic unit of trust, stating what an agent may do on its own and what needs a human. Autonomy is earned up a ladder from observe to operate, promoted by evaluation against fault scenarios, and enforced twice, inside the agent by policy awareness and outside it by Kubernetes RBAC, scoped service accounts and sandboxes.
A plugin registry for agent runtimes, added in one line, that packages what a generic model cannot know. Three plugins so far. analysis-facilities, with skills for HTCondor, JupyterLab, XCache, Rucio, ServiceX and Triton at UChicago, BNL and SLAC; atlas, with subagents and more than 25 skills for Rucio, AMI, pyhf, TRExFitter and TopCPToolkit; and hep-python-tools for good-practice Python. Together with the facility's managed settings and context files, it is the portable "playbook" layer of our agentic work.
usatlas/marketplace on GitHub →Every ATLAS job needs calibration and alignment data from the conditions database at CERN, at peaks above 20,000 requests a minute. With the ATLAS computing team we replaced the mesh of site-level Squid proxies with high-capacity regional Varnish caches behind location-aware DNS, and re-architected the Frontier launchpad on redundant Kubernetes clusters. The result is 30 percent faster delivery with a third of the servers, a sustained cache hit rate above 99 percent, and a year of uninterrupted operation. Metrics flow to Elasticsearch at UChicago, where an on-premise AI agent flags anomalous traffic.
The infrastructure we build and operate, and the testbeds where new patterns are proven.
A Kubernetes-based facility serving U.S. ATLAS physicists and their collaborators, with JupyterHub and BinderHub notebooks, Dask clusters, HTCondor batch that flocks onto the Midwest Tier2 Center, GPUs for machine learning and inference, and petabyte-scale Ceph storage with XCache data delivery. More than 850 registered users; an AI assistant and autonomous monitoring agents support users and operators.
af.uchicago.edu →Hosted HTCondor access points where a researcher signs in with a campus identity, joins a project and submits work to the Open Science Pool, campus clusters and collaboration resources. OSG Connect, launched in August 2013, was the first front door to the OSG for individual investigators and small campus labs and delivered 2.5 billion core hours to more than 650 projects. CMS Connect ran at UChicago for a decade; SPT Connect and PSD Connect, now on the Pile cluster and open to anyone in the Physical Sciences Division, run today.
Project page →A bare-metal Elasticsearch and Kibana platform at UChicago indexing 111 billion documents of operational data from ATLAS distributed computing, the Analysis Facility, worldwide perfSONAR network measurements and the conditions and data delivery systems. It serves ATLAS computing coordinators, shifters and facility operators, and now the AEGIS agents and the Analysis Facility assistant. Led by Ilija Vukotic; contact him to index or analyze operational data.
Project page →MANIAC Lab hosts the Scalable Systems Laboratory (SSL) of IRIS-HEP, the NSF Institute for Research and Innovation in Software for High Energy Physics. The SSL is the institute's testbed for new analysis systems, where ServiceX, Coffea, Dask and inference services are integrated at scale, benchmarked against the data rates of the High-Luminosity LHC, and taught in IRIS-HEP, HSF and ATLAS training events. The Lab leads the SSL and contributes to the IRIS-HEP Analysis Systems area and its agentic-analysis roadmap.
iris-hep.org/ssl →Built in the IRIS-HEP Scalable Systems Laboratory, binderhub.ssl-hep.org turns any Git repository into a reproducible JupyterLab session with guaranteed CPU, memory and GPU, pre-pulled images that cut ten-minute cold starts to seconds, Keycloak single sign-on with self-service groups for workshop leads, and multi-cluster federation that places sessions on the UChicago Analysis Facility, the SSL River cluster or the National Research Platform. It has served HSF-India, CODAS-HEP and PyHEP workshops, with Dask Gateway scale-out to HTCondor workers next.
binderhub.ssl-hep.org →Founded in 2005 by the University of Chicago and Indiana University, and joined in 2011 by the University of Illinois at Urbana-Champaign (Mark Neubauer, local PI) through the Illinois Campus Cluster, MWT2 is one of the largest Tier2 centers in the Worldwide LHC Computing Grid. It provides more than 40,000 CPU cores and 25 PB of storage for ATLAS simulation, reconstruction and analysis, and hosts the LOCALGROUPDISK and analytics services used by the U.S. ATLAS community.
More on our ATLAS work →RP1 is the incubator where new patterns are validated before they move into the production Analysis Facility. It is a multi-site Kubernetes platform spanning UChicago and Indiana University with a BinderHub front door, Keycloak identity, Dask Gateway and TaskVine scheduling, Kueue, and the ServiceX and Coffea analysis stack, all managed as declarative infrastructure with FluxCD.
rp1.hl-lhc.io →ServiceX, developed with IRIS-HEP, transforms experiment data into analysis-ready columnar arrays on demand and runs in production on the Analysis Facility. ServiceY is the Lab's next-generation delivery service, demonstrated at 300 Gbps on the FABRIC testbed, designed for the data rates of the High-Luminosity LHC.
ServiceX at IRIS-HEP →The Open Data Facility mirrors ATLAS Open Data in the United States and offers several ways to work with it, from BinderHub and JupyterHub notebooks to reproducible REANA workflows, all running on RP1. It supports education, outreach and AI-era research on open physics data, and is being developed with the CLARIPHY community.
The collaborations whose computing we operate or help build.
Completed projects whose ideas and software carry on in our current work.