Skip to content

Use Cases

End-to-end walkthroughs for representative scientific and data-readiness scenarios are located in use_cases/. Each follows one layout (Prerequisites, Setup, Execution as pasted prompts with expected results, Post-Conditions, Coverage, Cleanup) and covers data acquisition, code or skill registration, pipeline construction, and agent-driven execution. Some run on real scientific datasets; the skill-management demos use small fixture data.

Fixture data of a few megabytes is in the use case's own data/ folder in the repository; larger datasets are hosted in the DSAgt use-case data folder, and each Setup section gives the copy or download command.

DSAgt use cases

Use case Domain Summary
Microbial Isolates Genomics, short-read QC and assembly with fastp + megahit Register short-read QC and assembly codes, follow the genomics best-practice documents, and build a reproducible isolate-processing pipeline against real sequencing reads.
Cryo-EM Structural biology, EMPIAR-10017 β-galactosidase micrographs via CryoPPP DSAgt-assisted curation of cryo-EM data from the EMPIAR public archive (EMPIAR-10017 β-galactosidase micrographs via CryoPPP): register curation codes, ingest cryo-EM quality knowledge, and build a micrograph-preprocessing pipeline, with the AIDRIN AI-readiness check measuring the curation step before/after.
VASP DFT → AI-Ready Records Materials science, VASP DFT output to AI-ready records, via catalog skills and a registered code Convert VASP DFT output into AI-ready records: the agent discovers and installs a pymatgen skill from a catalog, authors a converter skill for a slab calculation, extends it to nudged-elastic-band calculations, registers that converter as a code, and runs it with provenance against a reference record (no DFT run, no HPC).
Tokamak Stability Fusion energy, M3D-C1 finite-element simulation data Register codes for reading, analyzing, and visualizing data from the M3D-C1 finite-element code, then use DSAgt to explore an example linear-MHD dataset, produce plots and spectra, repackage secondary data products as HDF5, and reconstruct the session as a script that reruns on other datasets.
AIDRIN AI data readiness, aidrin metrics (quality, fairness, privacy) on UCI Adult Apply AIDRIN through DSAgt to a single tabular dataset (UCI Adult), 15 metrics spanning data-quality, impact-on-AI, fairness-and-bias, and data-governance.
Genesis Skills for Data Curation Skill management, the external Genesis skill catalog driving a data-curation pipeline Sync the Genesis skill catalog, install the Croissant validation skill, ground the curation skills in the dataset's domain documents, and produce a datacard for a small curated dataset.
Plasma Turbulence Training Data (XGC) Plasma physics, gyrokinetic turbulence simulation (XGC) training-data prep Register the XGC preprocessing scripts as codes and drive them through DSAgt with provenance: check an ADIOS2 BP5 simulation case, summarize its physics content, preprocess it into GNN-ready npz files, validate the output, and reconstruct the pipeline. Advanced, HPC-scale data.
BlastNet → WELL Conversion Combustion CFD, BlastNet DNS trajectories to the WELL HDF5 format Develop a BlastNet-to-WELL converter from the format documents: author a conversion skill that carries the specifications as references, register the agent-written converter and a checker as codes, convert a sample trajectory with provenance, and compare the result with a reference WELL file held back from the agent, fixing the converter until the two match exactly.

Each use-case folder is laid out by role: README.md is the walkthrough; data/ holds fixture data small enough to keep in the repository; docs/ holds documents the agent reads; scripts/ holds code copied into the project; skills/ holds skills copied into the project; reference/ holds expected outputs and reference solutions that are not inputs to the demo. Only the folders a use case needs are present.

Dependencies follow one rule per kind. Python packages from PyPI that a walkthrough needs are an extra of dsagt named after the walkthrough (vasp-dft, combustion-simulation, tokamak-stability, plasma-turbulence), so pip install "dsagt[<use-case>] @ git+https://github.com/AI-ModCon/dsagt.git" installs them with dsagt (uv sync --all-extras in a checkout). Anything else (conda-only tools, libraries built from source) is installed by the use case's scripts/setup_env.sh into ~/dsagt-projects/.tools/<use-case>/, a shared tools directory beside the projects; the README's Prerequisites list what the script installs and what it needs already present. Larger data comes from the Google Drive folder linked above.

Adding a use case

Drop a README.md with frontmatter (title, domain, summary) into a use_cases/<name>/ folder; it is added to this table, its body is inlined as its own page, and it appears in the nav. A folder without frontmatter is left out. See hooks/gen_use_cases.py.