Skip to content

Architecture

DSAgt architecture

DSAgt provides a preconfigured agent platform with added capabilities for AI-ready scientific data processing and curation. The capabilities complement the agent platforms DSAgt runs on (Claude Code, Goose, Codex, opencode, Cline), which carry the model, the tool loop, and the native skills. Capabilities are exposed to the agent through one MCP server: skill discovery and installation, data-processing code execution with provenance, knowledge-base extension and retrieval, and explicit and cross-session memory. Observability through MLflow traces, episodic memory management, and vector-store indexing run in the server's background tasks. Every project also carries three base skills, skill-creator, datacard-generator, and aidrin, installed from their upstream repositories at init.

Capabilities

Code Registry (dsagt-server) The agent registers CLI data processing codes as skills (markdown files with YAML frontmatter whose frontmatter declares the executable) under <project>/skills/, beside the instruction skills. DSAgt handles dependency installation via uv run --with (uv installs with dsagt) and wraps every execution with dsagt-run for provenance capture. The agent discovers codes via search_registry.

Knowledge Base (dsagt-server) Hybrid semantic + BM25 search over ChromaDB collections partitioned by concern: code specs, the skill corpus, scientific documents, and per-project memory. A first dsagt init installs the built-in code specs and a default genesis skill corpus; the NeMo Curator reference collection and additional skill sources are also available (k-dense-ai, anthropic, antigravity, composio), and new collections can be added for the documents of a specific scientific pursuit. The agent searches via kb_search, ingests via kb_ingest, and saves user-confirmed facts via kb_remember. Opt-in episodic memory chunks and embeds each session turn into the per-project session_memory collection.

Provenance (dsagt-run) A wrapper around every registered-code execution. Records the command, arguments, exit code, duration, file counts, and truncated stderr to <project>/trace_archive/<record_id>.json and emits a code.execute span to the trace store. The server incrementally embeds those records into a code_use collection so past executions are retrievable, and the agent calls reconstruct_pipeline to render the trace archive as a reproducible workflow.

Observability (serverless MLflow) Traces are stored in an MLflow SQLite file at <project>/mlflow.db. DSAgt server actions are recorded there and the agent's LLM-call traces are translated from the on-disk session transcript to MLflow's Claude autolog trace format. View with dsagt traces <project>.

Skills Discovery DSAgt has MCP tools to connect to external GitHub skill repositories and search for skills for scientific workflows. The agent's platform discloses the skills installed in its native skills directory; DSAgt adds an extendable corpus of skills that can be searched and installed on demand, so an uninstalled skill takes no space in the agent's context. The agent searches via search_skills and installs via install_skill.

AI-Readiness Check (on by default) AIDRIN as the check around every tabular pipeline stage. aidrin is a registered code in every project, and with the check on the agent's instructions say the check for a tabular stage is the aidrin skill's quality baseline, run before and after each data transformation with the delta reported.

Memory DSAgt adds two memory extensions that complement the host agent's own memory (which condenses session facts into a managed set of Markdown files loaded into context). Explicit memory records user-confirmed facts as YAML (kb_remember / kb_get_memories). Episodic memory (opt-in) keeps a vector store of semantically chunked turn blocks and searches them by successive filtering (first to a session, then by regex over the query's key terms, then a vector ranking of what remains), so a long, multi-session history can augment agent context while the retrieval stays selective.

Project Layout

dsagt init prompts for the project location, defaulting to ~/dsagt-projects/<name>/ (enter /data/runs to place it at /data/runs/<name>/, or . for the current directory).

~/dsagt-projects/cheese-metagenome/
  .dsagt/                       # dsagt-internal state (hidden)
    config.yaml                 # project configuration (set by dsagt init)
    state.yaml                  # session log + memory cursor (owned by the MCP server)
    explicit_memories.yaml      # user-confirmed facts
  skills/                       # agent skills (SKILL.md + reference docs)
  trace_archive/                # code execution records (JSON, from dsagt-run)
  mlflow.db                     # serverless MLflow SQLite trace store
  kb_index/                     # knowledge base vector collections

  # Per-agent runtime config (one of, generated by dsagt init):
  #   claude:   CLAUDE.md, .mcp.json
  #   goose:    goose.yaml, .goosehints
  #   codex:    AGENTS.md, .codex-data/config.toml
  #   opencode: AGENTS.md, opencode.json
  #   cline:    .clinerules/, cline_mcp_settings.json (managed via cline mcp add)

Projects are registered in ~/dsagt-projects/projects.yaml, so dsagt <command> <name> works from any directory. The project's data (knowledge base, trace store, registered codes, skills, audit records) is agent-agnostic: re-running dsagt init for an existing project and choosing a different agent switches platforms and keeps everything accumulated (it prompts before any destructive change). dsagt rm <name> deletes the project directory and its registry entry.