# Praxist — Official Product and Documentation Corpus
> Praxist is an autonomous research product developed by Sapient Intelligence Pte Ltd. This file is generated from the website's canonical documentation, FAQ, and measured examples.
- Product: https://praxist.sapient.inc/en
- Company: https://sapient.inc
- Source: https://github.com/sapientinc/PRAXIST
- Product release date: 28 August 2026
## Product FAQ
### What is Praxist?
Praxist is an autonomous research system for measurable research problems that can be executed on a computer. It turns an already runnable project into a continuous, evidence-driven research run.
Across successive generations, parallel research agents develop candidate solutions; evaluators convert results into structured evidence; and a planning panel synthesizes that evidence into the research agenda for the next generation. The cycle continues until the search converges or the budget is exhausted.
You provide a runnable project and a measurable objective. Praxist orchestrates the research process that searches for the best-performing solution.
### How is Praxist different from manual tuning or AutoML?
AutoML tunes parameters within a predefined search space. Praxist runs the full research loop.
Parallel research agents can change methods, architectures, and strategies. Evidence from evaluation shapes the agenda for the next generation, while the Deep Innovation Gate (DIG) and Quality-Diversity (QD) allocation help the system escape local optima.
Praxist is closer to a self-directing research team than a search tool. If your researchers are already iterating on a problem manually, Praxist takes over the iteration loop itself.
### Is my project a good fit for Praxist?
Praxist delivers the most value when three conditions are met:
- **The objective is measurable:** there is at least one metric that meaningfully distinguishes better from worse, with a clear optimization direction.
- **The project already runs:** the baseline code, environment, and required data or simulator are in place and work without Praxist.
- **The best path forward is unknown.**
If a prerequisite is missing, Praxist stops and tells you exactly what is needed. It will not silently download unspecified datasets, invent a simulator, or fabricate baseline performance. That is a deliberate design principle.
### Do I need an API key, and what will it cost?
No API key is required in Codex-native mode; Praxist uses your authenticated Codex session. We also recommend using your own API key to access supported model APIs.
API costs are set by the provider and vary by model and usage. Total cost also depends on parallelism, the number of generations, and evaluation runtime. For cost-sensitive runs, start with a small representative workload before scaling up.
### How does Praxist protect my code and data?
Praxist provides three layers of protection:
- **Project isolation:** Praxist does not modify your original project. Run artifacts are stored separately.
- **Credentials:** API keys are entered through a masked local prompt and are not exposed in commands, shell history, or conversations.
- **Data collection:** Praxist does not collect data used in your experiments. It collects only limited system-level operational information, which you can disable at any time.
### How can I trust that a reported improvement is real?
Praxist uses three safeguards:
- **Preregistration:** Metrics, evaluation protocols, baselines, and acceptance thresholds are defined before the run.
- **Consistent evaluation:** Every candidate is measured through the same evaluator, and invalid or suspicious results are excluded.
- **End-to-end provenance:** Every reported improvement includes the evidence and lineage needed to inspect and reproduce it.
We recommend reviewing what the selected solution changed and testing it again in your own environment. Praxist's results are designed to be verifiable, and your own validation should be the final test.
### What if Praxist does not improve the result?
Praxist does not guarantee a specific metric improvement. It provides a rigorous research process and auditable evidence.
If a run does not meet its target, you still receive a negative-result evidence package, an audit report, and recommendations on whether to stop or redirect the research.
A negative result can still be valuable: it rules out tested approaches with evidence and helps prevent further investment in an unproductive direction.
### Is Praxist open source, and what terms apply to its outputs?
Praxist is licensed under the Fair Source License Agreement 1.0. The precise description is source-available: the complete source code is publicly available and may be viewed, downloaded, and modified. Subject to the license terms, Praxist may be used for internal business purposes and deployed within your own organization.
Organizations with aggregate annual revenue, including revenue from affiliates, below US$1 million may use Praxist commercially at no charge. Once annual revenue reaches or exceeds that threshold, the organization must contact the Licensor, Sapient Intelligence Pte Ltd, to negotiate a Commercial License.
The revenue threshold does not apply to qualifying teaching and academic research conducted by institutions of higher education, public research institutions, and nonprofit academic research organizations.
**Generated outputs:** no attribution is required for internal use. If an output is published externally or otherwise made available to third parties, the product-name attribution "Praxist by Sapient Intelligence" must be retained.
This FAQ is a summary only. If it conflicts with the Fair Source License Agreement 1.0, the terms of the license agreement control.
## Measured Examples
### Mobile crop-disease diagnosis
- Industry: Agriculture & Environmental Monitoring
- Result: ACCURACY 0.8909 → 0.90209 ↑
- Application: Use field photos to identify likely crop disease and route growers toward the right treatment, agronomist, or containment action.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/cassava-leaf-disease-classification/leaderboard.csv
### Classification from images plus engineered measurements
- Industry: Agriculture & Environmental Monitoring
- Result: LOG LOSS 0.10834 → 0 ↓
- Application: Combine visual evidence with measurements to classify parts, materials, crops, or defects when the labeled dataset is too small for image-only deep learning.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/leaf-classification/leaderboard.csv
### Deploy image models into new regions
- Industry: Agriculture & Environmental Monitoring
- Result: MACRO F1 0.108 → 0.40915 ↑
- Application: Train on established sites and retain accuracy at new branches, geographies, suppliers, or camera installations where backgrounds and class frequencies shift.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/iwildcam-2019-fgvc6/leaderboard.csv
### Visual identity matching with unknown detection
- Industry: Agriculture & Environmental Monitoring
- Result: MAP@5 0.32788 → 0.59833 ↑
- Application: Match people, animals, assets, products, or components to known identities from distinctive markings while explicitly handling previously unseen identities.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/whale-categorization-playground/leaderboard.csv
### Next-best-product recommendations
- Industry: Commerce, Marketing & Customer Experience
- Result: MAP@12 0.02177 → 0.03344 ↑
- Application: Recommend the most relevant products, offers, content, or replenishment items from behavioral history and catalog context.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/h-and-m-personalized-fashion-recommendations/leaderboard.csv
### Identify the words driving customer sentiment
- Industry: Commerce, Marketing & Customer Experience
- Result: JACCARD 0.71378 → 0.72392 ↑
- Application: Highlight the phrase behind praise, frustration, churn risk, or a complaint so teams see not only the score but the actionable reason.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/tweet-sentiment-extraction/leaderboard.csv
### Predict response or conversion from narrative plus context
- Industry: Commerce, Marketing & Customer Experience
- Result: AUROC 0.5996 → 0.8414 ↑
- Application: Estimate whether an application, appeal, campaign, lead, donation request, or support message will receive a positive response using both language and metadata.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/random-acts-of-pizza/leaderboard.csv
### Authorship, source, and style attribution
- Industry: Content, Media & Communications
- Result: LOG LOSS 0.41879 → 0.16327 ↓
- Application: Attribute documents or messages to likely sources, teams, content types, or style profiles for forensics, routing, brand consistency, and provenance analysis.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/spooky-author-identification/leaderboard.csv
### Rich visual attribute tagging
- Industry: Content, Media & Communications
- Result: MICRO F1 0.627 → 0.68815 ↑
- Application: Generate multiple structured attributes for products, creative assets, real estate, archives, or inspection photos to improve discovery and analytics.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/imet-2020-fgvc7/leaderboard.csv
### Legal Document Transcription
- Industry: Education & Public Services
- Result: F1 0.5985 → 0.97413 ↑
- Application: Transcribe difficult scanned records into searchable text for legal discovery, court and land-registry archives, document migration, and compliance workflows.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/kuzushiji-recognition/leaderboard.csv
### Consistent scoring of written submissions
- Industry: Education & Public Services
- Result: QUADRATIC KAPPA 0.82827 → 0.83838 ↑
- Application: Score essays, grant narratives, applications, audits, or quality reviews against a rubric to prioritize human attention and provide faster feedback.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/learning-agency-lab-automated-essay-scoring-2/leaderboard.csv
### Automated retinal screening and referral prioritization
- Industry: Healthcare & Life Sciences
- Result: QUADRATIC KAPPA 0.88891 → 0.93198 ↑
- Application: Triage screening images by disease severity so clinicians can focus first on patients most likely to need urgent follow-up while preserving human review.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/aptos2019-blindness-detection/leaderboard.csv
### Predict hidden biological or material traits from imaging
- Industry: Healthcare & Life Sciences
- Result: AUROC 0.52553 → 0.65882 ↑
- Application: Infer an expensive, invasive, or slow-to-measure property from multiple imaging channels, enabling earlier triage and more targeted confirmatory testing.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/leaderboard.csv
### RNA degradation and stability prediction
- Industry: Healthcare & Life Sciences
- Result: LOG LOSS 0.3631 → 0.22453 ↓
- Application: Predict how a biological or material sequence behaves at every position and under multiple conditions, helping researchers screen designs before costly experiments.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/stanford-covid-vaccine/leaderboard.csv
### Skin cancer detection from dermoscopy images
- Industry: Healthcare & Life Sciences
- Result: AUROC 0.9128 → 0.94608 ↑
- Application: Prioritize rare, high-risk cases from images and contextual metadata while explicitly optimizing sensitivity under severe class imbalance.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/siim-isic-melanoma-classification/leaderboard.csv
### Molecular and materials property prediction
- Industry: Industrial, Energy & Scientific R&D
- Result: MEAN-COLUMN RMSLE 0.06988 → 0.04969 ↓
- Application: Screen candidate materials or formulations virtually, reducing the number of simulations and physical experiments needed to find promising designs.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/nomad2018-predict-transparent-conductors/leaderboard.csv
### Remote asset-versus-hazard classification
- Industry: Industrial, Energy & Scientific R&D
- Result: LOG LOSS 0.20371 → 0.12707 ↓
- Application: Classify ambiguous targets in radar, sonar, thermal, or satellite imagery to protect offshore operations, shipping, borders, and remote infrastructure.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/statoil-iceberg-classifier-challenge/leaderboard.csv
### Real-time 3D hazard and asset detection
- Industry: Mobility, Logistics & Geospatial
- Result: MAP 0.042 → 0.15199 ↑
- Application: Build a perception layer for vehicles, robots, yards, mines, or warehouses that locates people, equipment, and obstacles in three dimensions so automated systems can act safely.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/3d-object-detection-for-autonomous-vehicles/leaderboard.csv
### Binary visual inspection and routing
- Industry: Operations, Risk & Forecasting
- Result: LOG LOSS 0.12216 → 0.00098 ↓
- Application: Classify an image into pass/fail, target/non-target, damaged/undamaged, or eligible/ineligible to automate a high-volume first decision.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/dogs-vs-cats-redux-kernels-edition/leaderboard.csv
### Operational case classification and routing
- Industry: Operations, Risk & Forecasting
- Result: ACCURACY 0.95342 → 0.96295 ↑
- Application: Assign each transaction, account, case, or asset to the right category from structured features for workflow routing and downstream decisions.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/tabular-playground-series-dec-2021/leaderboard.csv
### Recognize work steps and gestures from multimodal sensors
- Industry: Operations, Risk & Forecasting
- Result: LEVENSHTEIN 0.322 → 0.03374 ↓
- Application: Identify assembly steps, safety gestures, customer interactions, or operator actions by combining video, depth, and audio over time.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/multi-modal-gesture-recognition/leaderboard.csv
### Real-time abuse screening for digital channels
- Industry: Security, Trust & Compliance
- Result: AUROC 0.77842 → 0.95897 ↑
- Application: Score comments, chats, reviews, and community posts for likely abuse so platforms can warn users, prioritize moderation, or apply policy controls.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/detecting-insults-in-social-commentary/leaderboard.csv
### Policy-aware content moderation
- Industry: Security, Trust & Compliance
- Result: MEAN-COLUMN AUROC 0.98079 → 0.98803 ↑
- Application: Detect multiple policy violations in comments, reviews, chats, or cases so different behaviors can trigger different workflows rather than one blunt toxic/not-toxic rule.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/jigsaw-toxic-comment-classification-challenge/leaderboard.csv
### Automatic case, document, and ticket tagging
- Industry: Software, Knowledge & Productivity
- Result: MICRO F1 0.60685 → 0.79589 ↑
- Application: Assign multiple topics to support tickets, knowledge articles, legal documents, or incident reports for routing, search, analytics, and ownership.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/facebook-recruiting-iii-keyword-extraction/leaderboard.csv
### Context-aware text repair and missing-field recovery
- Industry: Software, Knowledge & Productivity
- Result: LEVENSHTEIN 5.55211 → 5.30562 ↓
- Application: Repair truncated messages, OCR gaps, transcription omissions, or incomplete product and case descriptions using surrounding context.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/billion-word-imputation/leaderboard.csv
### Multilingual customer and employee question answering
- Industry: Software, Knowledge & Productivity
- Result: WORD JACCARD 0.72756 → 0.88943 ↑
- Application: Answer questions from policies, help centers, product documentation, or case files in lower-resource languages without forcing users into English-only workflows.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/chaii-hindi-and-tamil-question-answering/leaderboard.csv
### Patent phrase similarity and deduplication
- Industry: Software, Knowledge & Productivity
- Result: PEARSON R 0.851 → 0.87388 ↑
- Application: Find equivalent or related technical concepts across patents, requirements, contracts, product catalogs, and knowledge bases despite different wording.
- Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/us-patent-phrase-to-phrase-matching/leaderboard.csv
## Technical Documentation
# Praxist
Praxist coordinates parallel research agents, experiments, evidence retention,
and synthesis across generations. It supplies the reusable research process;
your task project supplies the science.
Read the technical paper,
[*Praxist: From Experimental Artifacts to Solution Lineages*](https://arxiv.org/abs/2608.25955),
for the system design and evaluation.
## Research Infrastructure, Explicit Science
A coding agent is the recommended interface between the researcher and two
deliberately separate systems.
```mermaid
flowchart LR
HUMAN(("Researcher"))
AGENT(["Codex / Claude Code
recommended interface"])
PRAXIST["Praxist
generic research process"]
TASK[["Task project
scientific contract"]]
HUMAN --> AGENT
AGENT --> PRAXIST
AGENT --> TASK
PRAXIST <--> TASK
class HUMAN actor
class AGENT interface
class PRAXIST system
class TASK task
```
| Praxist owns | The task project owns |
|---|---|
| Peers, generations, orchestration, resource scheduling, and lifecycle | Objective, constraints, baseline, environment, and permitted changes |
| Agent runtimes, evidence transport, durable state, replay, and synthesis | Evaluator, metrics, protocol, evidence maturity, prompts, and roles |
The boundary keeps Praxist reusable across fields and makes the task project
the sole source of domain meaning. [Architecture](https://praxist.sapient.inc/en/docs/concepts/architecture)
defines the software boundary; [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects) defines
the scientific contract.
## Install Praxist
Install and configure Praxist with one command:
```bash
python3 -m pip install --index-url https://pypi.org/simple "praxist[agents,codex]" && praxist setup --interactive --install-skills codex
```
Or let an agent install from PyPI and follow the packaged setup runbook:
```text
codex --yolo
# or: claude --dangerously-skip-permissions
Install and configure Praxist using its packaged OOBE runbook. Stop after readiness checks.
```
Installation never selects a project or starts research. Before the separate
takeover step, read the [Quickstart](https://praxist.sapient.inc/en/docs/getting-started/quickstart) and
[Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task). The selected project must
already contain the code and local resources needed to run its baseline;
Praxist does not invent missing data, simulators, credentials, or measurements.
## Choose A Runtime Profile
- :material-rocket-launch: **Start without an API key**
---
Use an existing Codex login through Codex-native mode.
[Codex-native profile](https://praxist.sapient.inc/en/docs/getting-started/quickstart#codex-native-mode-no-api-key)
- :material-cash-multiple: **Run cost-efficient long research**
---
Prefer a high-cache-hit-rate [open-source model API](https://praxist.sapient.inc/en/docs/guides/open-source-model-apis)
after a representative quality and cache check.
[Model API selection](https://praxist.sapient.inc/en/docs/guides/open-source-model-apis)
## Continue After Setup
- :material-flask-outline: **Prepare a research task**
---
Define the research brief, prerequisites, and launch gates.
[Your first task](https://praxist.sapient.inc/en/docs/getting-started/first-task)
- :material-console: **Operate from the shell**
---
Use the direct CLI for lifecycle and monitoring operations.
[Direct CLI operations](https://praxist.sapient.inc/en/docs/guides/operators)
## The Research Brief Is the Control Surface
Takeover can inspect code and measure an existing baseline, but it cannot infer
the researcher's priorities. The brief should identify the objective,
credibility standard, allowed resources, exploration policy, and practical run
budget. Those decisions shape research direction, experiment throughput, and
retention from the first generation onward.
[Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task#write-the-research-brief) provides
a complete example and explains how takeover turns that brief into a validated
task project.
## Find a Specific Answer
| You want to... | Start here |
|---|---|
| Install and configure Praxist | [Installation](https://praxist.sapient.inc/en/docs/getting-started/installation) |
| Complete setup and hand off a project | [Quickstart](https://praxist.sapient.inc/en/docs/getting-started/quickstart) |
| Understand project prerequisites | [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task) |
| Use agent workflows | [Agent Skills](https://praxist.sapient.inc/en/docs/user-guide/skills) |
| Diagnose a failure or stall | [Troubleshooting](https://praxist.sapient.inc/en/docs/operations/troubleshooting) |
| Understand the research loop | [Research Loop](https://praxist.sapient.inc/en/docs/guides/research-loop-variant-generation-flow) |
| Configure a task harness | [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects) |
| Choose a scaffold or complete reference | [Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates) |
| Inspect a complete Python/JAX project | [Rocket Booster Recovery](https://praxist.sapient.inc/en/docs/examples/rocket-booster-recovery) |
| Inspect a complete native Rust project | [Rocket Booster Recovery (Rust)](https://praxist.sapient.inc/en/docs/examples/rocket-booster-recovery-rust) |
| Extend Praxist | [Developer Guide](https://praxist.sapient.inc/en/docs/guides/contributing) |
| Look up an exact command | [CLI Reference](https://praxist.sapient.inc/en/docs/reference/cli) |
The [Documentation Policy](https://praxist.sapient.inc/en/docs/about/documentation) identifies the sole owner of
each contract.
---
# Installation
Praxist release CI qualifies Linux on CPython 3.11 and 3.12. The package also
targets macOS and other CPython 3.11+ environments, but those combinations are
not continuously release-tested. Run `praxist doctor` on every host before
launching research. Install Praxist in the Python environment from which Codex
or Claude Code will operate. Task-specific packages, datasets, simulators, and
accelerator libraries remain in the task environment.
## Prerequisites
```bash
python3 --version
codex --version # when Codex is the operator interface
claude --version # when Claude Code is the operator interface
```
The selected agent interface must already be usable. Authentication requirements
for each runtime route are defined in [Credentials](https://praxist.sapient.inc/en/docs/guides/credentials).
## Install And Configure
The supported runtime extras install both maintained peer runtimes and the
Codex-native integration. Choose the agent application whose bundled skills
should be registered; each complete installation command is a single line.
```bash
# Codex
python3 -m pip install --index-url https://pypi.org/simple "praxist[agents,codex]" && praxist setup --interactive --install-skills codex
# Claude Code
python3 -m pip install --index-url https://pypi.org/simple "praxist[agents,codex]" && praxist setup --interactive --install-skills claude
```
Run the command in the intended active virtual environment when the host Python
is externally managed. `python3 -m pip` keeps the package and `praxist`
entrypoint tied to the same interpreter. The base `praxist` package supports
inspection and package-level CLI operations; the documented extras are the
complete research-runtime installation. The explicit public index prevents an
incomplete package mirror from silently omitting a pinned runtime SDK.
The command stops if installation, setup, or readiness fails. A successful
command also stops at that boundary: it does not select a project or launch
research. The
[Quickstart](https://praxist.sapient.inc/en/docs/getting-started/quickstart) is the sole description of the OOBE sequence and
its agent-managed alternative.
## What Installation Changes
The one-line flow:
- installs the tested agent-runtime dependencies;
- exposes `praxist` in the selected Python environment;
- registers bundled skills only for the selected agent host;
- writes only the configuration explicitly selected during setup;
- materializes writable complete examples outside the package;
- runs host diagnostics; and
- stops before project selection and research launch.
It does not install task training dependencies, CUDA, datasets, simulators,
human-facing Codex or Claude Code applications, or a collector service.
Codex-native support downloads a platform-specific runtime package of roughly
100-150 MB, so the first pip installation may take several minutes.
Same-name operator-owned skills are never replaced silently. Interactive setup
offers keep, backup-and-replace, or cancel; non-interactive setup reports the
conflict and stops.
Open the hosted documentation with:
```bash
praxist docs
```
Read the [Quickstart](https://praxist.sapient.inc/en/docs/getting-started/quickstart) and [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task)
before using the separate takeover command.
## Writable Complete Examples
Read-only package resources are never research workspaces. First-use setup copies
bundled complete examples to `${PRAXIST_EXAMPLES_HOME:-~/PraxistExamples}` and
prints each writable path. Existing destinations are preserved during upgrades.
The discovery and materialization commands, available projects, and asset
boundary are owned by [Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates).
## Verify
```bash
praxist --version
praxist doctor
praxist examples list
```
`doctor` checks each detected Praxist-managed skill host. Use
`praxist doctor --target codex` or `praxist doctor --target claude` to inspect
one host explicitly. If the shell cannot find `praxist`, activate the Python
environment used for installation or add that environment's script directory
to `PATH`.
## Uninstall
Stop active runs, remove Praxist-managed user state, then uninstall the Python
package from the same environment:
```bash
praxist uninstall --dry-run
praxist uninstall
python3 -m pip uninstall praxist
```
`praxist uninstall` removes only proven Praxist-managed skills, configuration,
local agreement/usage state, registry, and cache. `--keep-user-data` preserves
user records and caches. Pip removes the package from the active Python
environment. Research projects, writable examples, task environments, run
artifacts, agent applications, datasets, and task dependencies are never
removed.
## Source Development
Contributors use the repository environment rather than the operator install:
```bash
uv sync --group dev --extra docs
uv run praxist --help
```
See [Contributing](https://praxist.sapient.inc/en/docs/guides/contributing) for verification and
[Platform Support](https://praxist.sapient.inc/en/docs/operations/platform-support) for host boundaries.
---
# Quickstart
Praxist separates installation and configuration from research takeover. The
local-terminal lane keeps every setup prompt in the terminal. The agent-managed
lane lets Codex or Claude Code perform the same setup decisions in one
conversation. Both lanes write the same Praxist configuration and run the same
readiness checks; neither selects a project or launches research during
installation. They do not share an interaction controller or a second OOBE
state file.
Check [Installation](https://praxist.sapient.inc/en/docs/getting-started/installation) for host requirements and
[Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task) for project readiness before takeover.
## Local-Terminal Setup
Run the Codex or Claude Code one-line command from
[Installation](https://praxist.sapient.inc/en/docs/getting-started/installation#install-and-configure). It installs Praxist and
opens the local first-use wizard. To reopen only the wizard later, run
`praxist setup --interactive`.
An interactive terminal opens the first-use wizard automatically. Use Up/Down
and Enter to choose an item. Esc goes back or cancels without inventing a
choice. Setup covers these stages:
| Stage | Local interaction | Result |
|---|---|---|
| Install | Pip installs Praxist and its maintained runtime integrations into the selected Python environment. | `praxist` is available; no task-owned dependency is installed implicitly. |
| Legal terms | Review the Fair Source License, User Agreement, and data notice in a temporary scroll view, then explicitly agree or cancel. | Acceptance records the exact legal bundle version and digest; it does not enable optional data collection. |
| Privacy | When collection is available, separately choose whether to share pseudonymized product-usage status. Nothing is preselected. | Existing consent is preserved; cancellation leaves consent unset. |
| Runtime | Choose one setup profile combining an API provider, agent runtime, concrete model, and authentication mode. | A coherent profile is written to the user configuration. |
| Readiness | Register skills for the selected agent host, materialize writable examples, and run host diagnostics. | Installation finishes without selecting a project or starting a run. |
API keys are entered only in the local terminal. Praxist displays one `*` for
each character, supports paste and backspace, and never places the raw value in
the command line, shell history, agent conversation, or task project. Esc
cancels key entry.
## Agent-Managed Setup
Start Codex or Claude Code:
```bash
codex --yolo
# or
claude --dangerously-skip-permissions
```
Then ask:
```text
Install and configure Praxist. Follow the packaged OOBE runbook and stop after
readiness checks. Do not select a project or start research.
```
The agent reads `docs/agents/oobe-install.md` and uses its structured choices
for legal acceptance, privacy, and runtime decisions. It links to the scrollable
legal documents and must not infer acceptance or accept on the operator's
behalf. An API provider key is never requested in chat. When an API-backed setup
profile is selected, the agent pauses for the operator to enter the key through
Praxist's local masked prompt.
Pip installation is only the package boundary; it does not imply that the
operator accepted legal terms or selected a runtime. The agent must continue
the OOBE runbook rather than report that setup is complete. Its first state
query is:
```bash
praxist setup --agent-managed
```
Its JSON identifies the next required decision. The agent reruns it after each
decision and cannot claim setup decisions are complete until
`setup_decisions_complete` is `true`. Existing credentials, API provider defaults,
and doctor readiness never count as an operator profile selection.
This lane and the local-terminal lane share only configuration, validation, and
recovery contracts. The agent does not emulate terminal keystrokes, and the
local wizard does not reproduce the agent workflow.
## Start Research After Reading the Manual
Before handing over a project, read [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task). It defines
what the project must already provide, what Praxist will add, and which choices
the takeover prompt must communicate. Installation success alone is not launch
authorization.
When the project is ready, start takeover as a separate command:
```bash
# Codex
praxist --takeover --task-path /absolute/path/to/research-project
# Claude Code
praxist --takeover --operator claude --task-path /absolute/path/to/research-project
```
## Choose a Setup Profile
The wizard presents coherent setup profiles that bundle an API provider, agent
runtime, concrete model, and authentication method.
### Codex-Native Mode: No API Key
Choose Codex-native mode to use the existing saved Codex login. This is the
shortest first-use path and does not write an API provider key. Only this explicit
profile verifies ChatGPT subscription authentication. If the SDK-pinned Codex
currently uses an API-key login, setup opens its local interactive login flow
and verifies the ChatGPT login before readiness checks continue. Other
profiles neither require nor inspect this login.
### API-Backed Profiles: Long Runs
For sustained, cost-sensitive research, prefer an
[open-source model API](https://praxist.sapient.inc/en/docs/guides/open-source-model-apis) whose cache reuse,
quality, and throughput have been checked on a representative workload. The
selector includes maintained DeepSeek API, OpenRouter API, and Anthropic API
profiles. Enter the selected provider's key at the local masked prompt when
requested.
Other supported API-backed setup profiles remain available in the same selector. To inspect
the exact current profile contract without changing configuration:
```bash
praxist setup --list-profiles
```
## Reopen a Step
The OOBE does not create a separate completion marker. It derives current state
from the legal-terms acceptance record, existing Praxist configuration,
the profile ID recorded by an explicit setup selection, separate product-usage
consent, diagnostics, and task artifacts. A repeated setup keeps the current
recognized profile selected and never launches takeover. An existing key for
the selected API-backed setup profile is preserved without a second prompt.
Reopen only the step you need:
```bash
praxist setup --interactive
praxist --takeover
praxist --takeover --operator claude
```
Review or verify the License and User Agreement independently with:
```bash
praxist user-agreement review
praxist user-agreement status --json
```
Use an explicit project path when the project is not the current directory:
```bash
praxist --takeover --task-path /absolute/path/to/research-project
```
Explicit `setup --profile` and provider automation flags remain available for
controlled non-interactive provisioning. They do not infer legal acceptance or
an operator profile choice.
## Bundled Complete Examples
Installation creates writable Python/JAX and Rust reference projects under
`~/PraxistExamples` and prints their absolute paths. Praxist never runs against
or writes into the read-only copies in its source or package directory.
Existing working copies are preserved during upgrades.
[Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates) owns the available
projects, copy commands, and guidance for choosing a starting point.
## Check Progress
Ask Codex or Claude Code in natural language:
```text
Report current research progress and list the strongest variant in every
completed generation with its task-defined performance metrics.
```
For direct shell operation:
```bash
praxist status --json
praxist --monitor --latest
```
`Ctrl-C` closes only the foreground monitor. It does not stop the research run.
Ask the agent to stop the current run, or use `praxist stop ` directly.
## Next
- [Installation](https://praxist.sapient.inc/en/docs/getting-started/installation) defines package and platform boundaries.
- [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task) explains project prerequisites and takeover.
- [Product Usage Controls](https://praxist.sapient.inc/en/docs/operations/product-usage) explains consent commands.
- [`LICENSE.md`](https://github.com/sapientinc/praxist/blob/main/LICENSE.md) is the canonical software license.
- [Praxist User Agreement](https://praxist.sapient.inc/en/docs/legal/user-agreement) defines the service terms accepted with it.
- [Direct CLI Operations](https://praxist.sapient.inc/en/docs/guides/operators) is the shell lifecycle guide.
---
# Your First Task
Praxist adapts an **existing runnable research project** into a task project.
The original project supplies the executable baseline; the task project tells
the generic framework what may change, how evidence is measured, and what
counts as credible progress.
```mermaid
flowchart LR
PROJECT(["Runnable project
code + environment"])
TASK[["Task project
scientific contract"]]
PRAXIST["Praxist
generic research process"]
EVIDENCE[("Task-local run
evidence + reports")]
PROJECT --> TASK --> PRAXIST --> EVIDENCE
class PROJECT source
class TASK task
class PRAXIST system
class EVIDENCE artifact
```
## What Must Exist Before Takeover
| Required input | Ready means |
|---|---|
| Research code | The baseline implementation and normal entrypoint are present. |
| Runtime | An existing interpreter, environment, container, or remote path can import the required dependencies. |
| Data or simulator | Every required asset is locally reachable through the project's normal interface. |
| Baseline path | Training, optimization, simulation, inference, or evaluation runs without Praxist. |
| Measurable objective | At least one metric distinguishes candidates and its direction is known. |
| Scientific constraints | Forbidden changes, validity conditions, and important tradeoffs can be stated. |
Prior results and technical documents improve initialization but are not always
required. Takeover reports missing prerequisites instead of downloading an
unknown dataset, inventing a simulator, or fabricating baseline performance.
## Write the Research Brief
The brief is the operator's main scientific input. It should settle:
1. **Objective:** what should improve and which tradeoffs matter?
2. **Evidence:** which metrics, protocol, and maturity level make a result
credible?
3. **Execution:** which environment and local assets are allowed, and what is
the realistic compute budget?
4. **Exploration:** should literature lookup, the
[Deep Innovation Gate (DIG)](https://praxist.sapient.inc/en/docs/guides/deep-innovation-gate),
[Quality-Diversity (QD)](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator), and
constructive-peer guidance be active?
5. **Operation:** how many peers and generations are appropriate, and may a
validated run launch unattended?
One or two bounded calibration runs often reveal where a long-run brief needs
revision. Readiness and scientific-integrity gates still apply; a brief cannot
authorize Praxist to bypass an unresolved contract.
??? example "Detailed takeover brief"
Adapt the intent to the project rather than copying the numbers.
```text
Invoke `praxist-takeover` in Codex or Claude Code.
Use the Praxist checkout at "/path/to/Praxist" on branch "main". Treat the
current directory as the existing research project. Verify that its current
environment can run the unchanged accelerator-backed training and evaluation
path. If no compatible environment exists but every required dependency is
locally available, create an isolated task environment without changing the
system Python.
Create a separate Praxist task project. Configure 12 peers, 30 generations,
and a generation duration justified by measured baseline runtime. Disable
public literature lookup. Enable QD, enable DIG only for absolute generation
zero, and use the constructive-peer ratio as a soft target.
If baseline performance is missing and the project already contains everything
required to measure it, run a bounded baseline benchmark and record metrics
with their provenance. Use the Praxist agent runtime, API provider, and model
selected during setup.
Define every metric direction, evidence maturity requirement, and
protocol-integrity check. Configure durable parent lanes to retain credible
Pareto-optimal solutions across genuinely different metric dimensions. Keep
partial, diagnostic, suspect, and protocol-failed evidence visible as
follow-up signals without treating it as clean parent evidence.
Do not download a new dataset or replace the project's existing simulator or
runtime. After mandatory evaluator, lane-routing, readiness, and runtime gates
pass, launch in detached mode without optional follow-up questions. Report the
task path, evidence contract, lane rules, generation-close policy, run ID, and
monitor command.
```
## Start Takeover
From the shell:
```bash
praxist --takeover --task-path /absolute/path/to/research-project
```
Codex is the default operator interface; add `--operator claude` for Claude
Code. In an existing agent conversation, invoke `$praxist-takeover` in Codex or
`/praxist-takeover` in Claude Code. Use `praxist-takeover-codex` only for the
explicit no-key Codex-native profile.
Use `praxist-task-initialization` to create or repair the harness without
launching, or `praxist-interactive-task-init` for confirmation-first design.
[Agent Skills](https://praxist.sapient.inc/en/docs/user-guide/skills) owns the complete goal-to-skill map.
## What Praxist Adds
Initialization adds the smallest practical harness around existing assets:
| Harness area | Purpose |
|---|---|
| Task contract | Objective, scope, permitted changes, metrics, evidence policy, roles, and run settings |
| Evaluator and baseline record | One reproducible path from a candidate to structured metrics with provenance |
| Retention and close policy | Reachable durable, Pareto, diagnostic, parent, maturity, and generation-boundary decisions |
| Resource observation | Unchanged-baseline timing and bottleneck evidence used to plan concurrency |
| Task tests | Evaluator output, lane reachability, maturity, resource handoff, and launch readiness |
| `experiments/` | Run artifacts outside Praxist source and stable project code |
The exact schema and precedence rules live only in
[Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects).
## Validate Without Starting
Takeover performs validation automatically. For direct inspection:
```bash
praxist resolve /absolute/path/to/task
praxist doctor --task-path /absolute/path/to/task
```
When the task requires ratios, validate a real evaluator summary through the
same serializer used in production:
```bash
praxist resolve /absolute/path/to/task \
--result-summary /absolute/path/to/evaluation_summary.json
```
## What Happens During Takeover
```mermaid
flowchart LR
DISCOVER(["Discover
project / runtime / baseline"])
DESIGN(["Design
metrics / evidence / roles / resources"])
VERIFY(["Verify
task tests / resolve / doctor"])
LAUNCH(["Launch
detached run / status / monitor"])
DISCOVER --> DESIGN --> VERIFY --> LAUNCH
class DISCOVER,DESIGN,VERIFY,LAUNCH phase
```
The operator agent:
1. identifies the active Praxist installation, project, execution environment,
local assets, technical context, and prior evidence;
2. reuses measured baseline evidence or offers a bounded measurement when every
prerequisite is available;
3. turns the brief into task-owned metrics, ranking, protocol integrity,
maturity, retention, close, role, prompt, and exploration contracts;
4. observes the unchanged baseline execution path to estimate runtime and the
actual resource bottleneck without changing its backend;
5. creates the harness by referencing existing assets instead of cloning the
project into Praxist;
6. runs task, evaluator, lane-routing, resource, resolve, and runtime checks;
7. launches only after mandatory gates pass, then reports the task path, run ID,
lifecycle state, and monitor command.
If a prerequisite or scientific decision is genuinely unresolved, takeover
stops at that point and names the missing input. It does not weaken evaluation
to make launch succeed.
---
# Agent Skills
Praxist skills are operator workflows for Codex and Claude Code. Invoke a skill
as `$name` in Codex or `/name` in Claude Code. Skill files instruct the agent;
they do not run as background services and they do not replace the `praxist`
CLI.
## Choose by Goal
| Goal | Codex | Claude Code |
|---|---|---|
| Learn Praxist and inspect host readiness | `$praxist-onboarding` | `/praxist-onboarding` |
| Install or repair Praxist runtime dependencies | `$praxist-runtime-install` | `/praxist-runtime-install` |
| Build or repair a task harness without launching | `$praxist-task-initialization` | `/praxist-task-initialization` |
| Confirm task design interactively | `$praxist-interactive-task-init` | `/praxist-interactive-task-init` |
| Initialize and launch with a configured API provider | `$praxist-takeover` | `/praxist-takeover` |
| Initialize and launch with a saved Codex login | `$praxist-takeover-codex` | `/praxist-takeover-codex` |
| Start, stop, resume, monitor, or inspect runs | `$praxist-control` | `/praxist-control` |
| Diagnose run health or generate reports | `$praxist-diagnostic` | `/praxist-diagnostic` |
| Gather literature and benchmark context | `$praxist-scientific-research` | `/praxist-scientific-research` |
| Draw a terminal line chart | `$terminal-line-plot` | `/terminal-line-plot` |
The complete catalog and activation descriptions are generated from the actual
skill metadata in [Skills Reference](https://praxist.sapient.inc/en/docs/reference/skills).
After installation, read [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task).
Then invoke the matching takeover explicitly without having to remember a skill
name:
```bash
praxist --takeover --task-path /absolute/path/to/research-project
```
For agent-managed installation, the agent follows
[the OOBE runbook](https://praxist.sapient.inc/en/docs/agents/oobe-install) and stops after readiness checks.
Takeover remains a separate post-manual action.
## Install or Refresh
An operator installation registers the bundled skills automatically. Refresh
them after upgrading Praxist:
```bash
praxist install-skills --target codex --replace
praxist install-skills --target claude --replace
```
Run only the command for the host you use. Codex installs under
`${CODEX_SKILLS_DIR:-~/.agents/skills}`; Claude Code installs under
`${CLAUDE_SKILLS_DIR:-~/.claude/skills}`.
This replaces only same-name Praxist skills managed by the current package.
Unrelated user skills are not touched. An operator-owned same-name path is
reported as a conflict and remains unchanged. Interactive setup can preserve a
backup before an explicitly approved replacement.
Source contributors can use symlinks for fast iteration:
```bash
bash scripts/install_codex_skills.sh
bash scripts/install_codex_skills.sh --target claude
```
## Workflow Boundaries
- Onboarding inspects and explains; it does not launch research.
- Task initialization writes only the selected task project.
- Control owns lifecycle operations and uses the CLI.
- Diagnostic is analysis-only unless the user explicitly asks for task-level
improvement.
- Scientific research records sourced context; it never claims literature as
measured task performance.
Task construction and takeover are explained in
[Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task). Report semantics live in
[Run Reports](https://praxist.sapient.inc/en/docs/guides/user-facing-reports-and-init), and lifecycle commands
live in [Direct CLI Operations](https://praxist.sapient.inc/en/docs/guides/operators).
---
# Task Projects
Task projects are research problems that Praxist runs. They are explicit inputs, not
bundled system plugins.
```mermaid
flowchart LR
PROJECT(["Research project
code / environment / assets"])
TASK[["Task project
objective / evaluator / task.yaml"]]
PRAXIST["Praxist
generic orchestration"]
ARTIFACTS[("experiments/
task-local run artifacts")]
PROJECT --> TASK --> PRAXIST --> ARTIFACTS
class PROJECT source
class TASK task
class PRAXIST system
class ARTIFACTS artifact
```
This page is the sole detailed owner of the task-project contract. Tutorials
summarize the workflow and link here rather than defining a second schema.
Planning uses Principal Investigator (PI) agents; multi-PI topologies add a
Chair that consolidates their proposals. [Panel Topology
Prompts](https://praxist.sapient.inc/en/docs/concepts/panel_topology_prompts) owns those roles.
## Location
For local dogfood, put task projects under the ignored repository-root
`tasks/` directory:
```bash
tasks/my_research_task/
```
For real collaboration, keep the task in its own Git repository and pass its path:
```bash
cd /path/to/task-project
praxist start --daemonize --json
# or pass --task-path explicitly from another directory
```
## Required Shape
A task project normally contains:
- `task.yaml` for task identity, workflow selection, metric, plugin refs, and
execution defaults;
- `description.md` for stable task context;
- `roles/` for task-local role skills. Praxist resolves the declared peer
RoleSkills before the research loop, injects each peer's agenda-assigned
Markdown contract into that peer's prompt, and records the effective role
reference plus content hash on the runtime request;
- `audit_rules/` for task-local proposal, result, or agenda criteria; these
should be declarative YAML or Markdown by default, not Python framework hooks;
- `evaluations/` for task-local evaluation logic. Expensive executable tasks
should expose one public evaluator command, normally
`evaluations//run.py`, so agents do not call internal harness scripts
directly;
- `assets/` for harness code, optional reference implementations, fixtures,
data metadata, and literature packs;
- task-local tests for the harness and any optional reference implementations.
Hardware planning observes the unchanged baseline and records its actual
backend. A task declares only profiles that its public evaluator can use; it
does not infer accelerator requirements from host inventory. Process handoff,
NVIDIA/CUDA UUID rules, natural-unit parallelism, and supply feedback are owned
by [Central Experiment Scheduler](https://praxist.sapient.inc/en/docs/guides/central-resource-scheduler).
Use `templates/tasks/template` as replaceable scaffolding. Complete examples
serve a different purpose and must run from their installed writable copies.
[Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates) owns that distinction and
the available starting points.
## `task.yaml` Contract
The task descriptor is the only machine-readable source Praxist needs from a task
project at startup. A typical descriptor includes:
- task identity: stable id, name, version, and description path;
- workflow defaults: enabled stages, generation count, cohort size, and run
defaults;
- metric contract: primary metric name, direction, optional secondary metrics,
mature-evidence ratios, constructive target, and cooperative launch guard;
- plugin refs: generic workflow stage, agent runtime, API provider, tool, budget, panel,
and graph refs;
- task-local refs: roles, audit rules, evaluations, and budget profiles;
- assets: harness paths, baseline files, literature packs, dataset metadata, and
optional reference implementations;
- task entrypoints: the public evaluation command and any optional internal
harness runners;
- output policy: default experiments directory and artifact retention notes.
- optional runtime environment: task execution cwd, venv/python path, PATH
additions, and non-secret task env vars.
- agent reasoning policy: `agent.reasoning_effort` applies one `auto`, `off`,
`low`, `high`, or `max` choice across API providers to peers and planning
calls;
omitted values default to `max`. Runtime-specific mappings are defined once in
[Agent Runtimes](https://praxist.sapient.inc/en/docs/guides/agent-runtimes#reasoning-policy).
- generation-scoped Deep Innovation Gate (DIG) and Quality-Diversity (QD)
policy: independent enable switches, generation scope, candidate-pool
requirements, diversity-cell fields, allocation caps, and task-owned
allowed/disallowed file rules.
- Gems policy: `gems.enabled`, optional reset cadence, compact Gem caps when
reset is enabled, relevant lanes, metric keys, and task-owned maturity
thresholds for staged or full-coverage evaluation protocols.
- Frontier lanes: optional `evaluation.frontier_lanes` entries that keep
mature candidates, promising validation candidates, and diagnostics in
separate task-owned evidence streams. Without frontier lanes, Praxist keeps the
legacy primary-metric frontier but does not materialize an incubator-style
lane for preliminary, aligned, or partial signals. Every lane should declare
`parent_eligible`: true only for mature durable parent lanes, false for
lower-stage and diagnostic lanes.
- Research-loop tools: declare the selected `tool_server:*` refs so resolve and
runtime agree. [Tool Servers](https://praxist.sapient.inc/en/docs/guides/tool-servers) owns the catalog,
[Scientific Literature Lookup](https://praxist.sapient.inc/en/docs/guides/scientific-literature-lookup) owns source
policy, and [Run Reports](https://praxist.sapient.inc/en/docs/guides/user-facing-reports-and-init) owns report
behavior.
- Metric directions: every metric used for frontier ordering, baseline-beat
detection, dimension winners, or charts must resolve to an explicit task-owned
`maximize` or `minimize` declaration. Result aliases inherit the direction of
their configured source metric. Unknown direction remains unknown; reports do
not guess that it should be maximized.
- QD plan versus evidence: `planned_dimensions` describes allocation intent;
`design_dimensions` describes the implementation that actually ran. Praxist
never fills missing evidence from the plan. The complete behavior is defined
in [Quality-Diversity Allocation](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator).
The descriptor may use the `praxist_plugins` block to bind generic plugin
refs and task-local component refs. Do not use task descriptors to smuggle
system code into core.
## Baseline Records
Serious task projects should keep compact baseline evidence under
`assets/baselines/`. Prefer:
- `results.jsonl` for machine-readable metric rows;
- `curated_baseline_summary.md` for human-readable interpretation and
provenance;
- `baseline_performance_status.md` for measurement status, command, data
source, environment, and any missing requirements.
If task initialization can measure the baseline on the current machine but no
baseline record exists, Codex should ask the operator whether to write explicit
zero placeholders or run a task-local baseline benchmark first. A benchmark run
belongs under `experiments/baseline_bench_/`, should use the public
task evaluator or documented baseline command, and may use bounded parallelism
only when the task and hardware make that safe. Zero placeholders must be marked
as placeholders, not measured performance facts.
Files under `assets/baselines/` preserve evidence and provenance; they do not
implicitly configure runtime comparisons. Verified values must also be declared
under `task.yaml:baselines` with their metric direction. `praxist resolve` and
`praxist start` warn when a conventional parseable result asset is present but
that declaration is empty. The warning is advisory and never auto-imports or
trusts task files.
## Evaluator Launch Readiness
Task initialization proves an evaluator before expensive fan-out. It first
exercises the task-appropriate build/load/startup boundary in the actual runtime,
validates the evaluator's public CLI, function, RPC, simulator, container,
notebook, or service contract as applicable, and then runs a **one-unit canary**
through the public evaluator, the central scheduler when Praxist owns launch,
and the canonical summary writer. "One
unit" is deliberately task-defined: it is the smallest valid case for that
task, not a framework-wide seed, epoch, split, iteration, or hardware rule.
The resulting summary must validate and project into a finding before wider
execution begins. Any implementation or command change requires a new canary.
The canary distinguishes broken execution from a valid weak or negative
scientific result; it does not establish performance or mature evidence. A task
that explicitly claims independently trusted evaluation must additionally
provide a task-owned verifier and demonstrate that peers cannot replace the
authoritative result and that altered or unattested evidence is rejected.
Peer-authored evaluators retain the normal path and do not inherit an external
attestation requirement.
## Override Rules
Operators can override run settings from CLI or config. Overrides should change
runtime choices, model profiles, budget envelopes, generation count, or local
paths. They should not mutate the task source.
The intended priority is:
```text
CLI args > explicit env vars > override spec > task.yaml defaults
```
Credentials are the exception: raw secrets are resolved by Python credential
resolution and are never copied into `task.yaml`.
## Experiments Directory
Task run outputs belong under a task-local ignored directory, normally:
```text
/experiments/
```
Do not write long-run task outputs into `templates/`, `examples/`, or `praxist/`.
Selection order is explicit `--run-dir`, then `$RUN_DIR`, then a timestamped run
under `/experiments/`. Praxist rejects destinations inside its own
source checkout. Pass `--run-dir` explicitly for an external output root.
The detached launcher log is `/logs/launcher.nohup.log`.
Task harnesses publish structured result summaries and findings. They do not
write frontier, Gems, prompt, report, or memory state directly. Every ranked
metric declares its own direction, and negative evidence uses structured
valence/failure metadata rather than prose alone. The canonical, validation,
derived, audit, and partial artifact roles are defined once in
[Architecture](https://praxist.sapient.inc/en/docs/concepts/architecture#state-and-replay).
## Runtime Environment
Tasks that need a specific virtual environment or executable can declare it in
`task.yaml`:
```yaml
runtime_environment:
cwd: task_project # task_project, run_dir, or a task-relative path
venv: .venv # task-relative or absolute path
# python: .venv/bin/python # optional override; inferred from venv otherwise
path_prepend:
- bin
env:
TASK_MODE: dogfood # non-secret task vars only
```
At startup Praxist validates the configured paths by default, injects
`PRAXIST_TASK_VENV`, `VIRTUAL_ENV`, `PRAXIST_TASK_PYTHON`, `PRAXIST_TASK_SHELL_PREFIX`,
and prepends the venv/python directories to `PATH` for agent runtime sessions.
Task prompts and harness commands should prefer `$PRAXIST_TASK_PYTHON` when present
and otherwise fall back to `python`.
When a task interpreter is declared, experiment children do not inherit the
Praxist runner's `PYTHONPATH` or `PYTHONHOME`. A task that genuinely requires
either variable must declare it in `runtime_environment.env`; Praxist then
treats that value as task-owned. This boundary prevents an older task Python
from importing packages out of the runner environment while retaining explicit
task-specific import layouts.
Do not put raw API keys in `runtime_environment.env`; model and tool secrets
belong to credential resolution.
Keep dependency installation as a separate task setup step. Task-specific
non-secret variables, interpreter selection, and import paths belong in
`runtime_environment`; start and resume remain owned by the Praxist CLI.
## Protocol Intent Belongs To The Task Owner
Praxist does not impose a universal full-protocol-only policy. Task
initialization should resolve one protocol-intent table from the current user
instruction first, then compatible project evidence, then a proposed default.
For every evaluator mode, the table states whether that mode may launch, rank,
count as mature, supply a durable parent, and satisfy close.
An explicitly requested partial, scout, reduced-coverage, or otherwise
incomplete protocol is valid. Its summaries must still report the actual stage,
effort, coverage, and integrity facts, and the task's maturity, lanes, Gems, and
close settings must all encode the same choice. Praxist should preserve useful
signals from other modes without presenting undeclared deviations as mature
evidence. Launch validation must inspect structured evaluator modes and output
metadata, never reject commands because their text happens to contain words
such as `smoke`, `scout`, or `partial`.
The configurations below are recommended defaults when the user and project do
not specify different semantics. They are not global restrictions.
## Deep Innovation Gate, Quality-Diversity, And Gems Defaults
For newly initialized real research tasks, the current recommended default is
continuous evolution with DIG limited to absolute gen0, independent QD enabled
for both the gen0 DIG pool and later PI synthesis, and periodic Gems reset
disabled:
```yaml
generation_policy:
max_generations: 8 # or another value suitable for a real research run
cohort_size: 5
per_generation_hours: 5
evaluation:
maturity_policy:
min_effort_ratio: 0.75
min_coverage_ratio: 0.80
require_ratio_gate: true
complete_stage_labels: [complete]
preliminary_stage_labels: [preliminary, aligned]
constructive_peer_mix_enabled: true
constructive_target_ratio: 0.75
launch_guard:
enabled: true
# Observed p90 runtimes of ordinary heavy work and the evaluator whose
# evidence is authorized for normal close.
estimated_heavy_eval_minutes: 0
estimated_close_grade_eval_minutes: 0
safety_factor: 1.25
synthesis_trigger:
mature_quorum_fraction: 0.25
quality_diversity:
enabled: true
initial_generation_enabled: true
later_generations_enabled: true
max_same_diversity_cell_peers: 1
max_same_mechanism_family_fraction: 0.34
max_same_intervention_surface_fraction: 0.50
gems:
enabled: false
selection_policy: mature_evidence_top_k
# Smoke placeholder; real tasks use their complete-protocol unit count.
min_mature_eval_units: 1
evidence_stage_min_units:
complete: 1
max_resets: 3
max_gems_per_reset: 4
max_gems_total: 4
max_gems_per_family: 2
prompt_max_gems: 4
archive_ordinary_findings: true
```
The excerpt keeps operator-facing evaluation, QD, and Gems controls visible.
Task initialization manages DIG's internal planner settings from the requested
enablement, generation scope, and runtime budget.
For tasks with task-defined mature/complete evidence, use a positive mature
quorum so raw progress or diagnostic findings cannot become normal completion.
`mature_supply_fraction` only prioritizes evidence production; it is not a
close gate. Use `0.0` only when the task intentionally has no separate
close-grade evidence contract and the operator explicitly accepts
information-density closing.
When mature/complete evidence is required for normal close, the finalized task
must satisfy:
```text
estimated_close_grade_eval_minutes * safety_factor
< effective_generation_close_horizon_minutes - drain_margin_minutes
```
Use the close-authorized evaluator's observed p90 runtime and at least a
30-minute drain margin unless measured publication/shutdown latency requires
more. `estimated_heavy_eval_minutes` may separately describe a longer optional
protocol; older tasks that omit the close-grade field use the heavy estimate as
a compatibility fallback. The effective horizon is the earliest enabled
generation or synthesis hard bound, including an enabled adaptive ceiling.
Praxist rejects a declared required-evidence contract that cannot finish by
construction. A user-authorized reduced or late-signal protocol remains valid
when its maturity, lane, launch, and close settings consistently describe that
intent.
Smoke fixtures may keep `max_generations: 1` while still declaring these blocks
so startup and config parsing stay visible; they may set later-generation QD
and next-generation constructive feedback to `false` because neither can take
effect in a one-generation run. Enable periodic Gems reset only
after an operator request or a diagnostic pass identifies a performance ceiling
and recommends a reset cadence. `reset_interval_generations` is meaningful only
when `gems.enabled: true`.
When the user-approved maturity contract uses ratios, each canonical evaluator
summary must emit `effort_ratio` and `coverage_ratio` in a supported scalar fact container such
as the summary root, `metrics`, `extra`, or `current_aggregate`. `effort_ratio`
is actual training/search/optimization effort divided by the task-defined
mature reference effort. `coverage_ratio` is completed required evaluation
units divided by total required units. Praxist uses one maturity extractor for the
source summary and its auto-materialized finding, so task code must not rewrite
the same facts into a second artifact. A standalone manually authored result
finding with no canonical summary reference must carry the ratios itself.
Task-specific stage labels remain audit context. Use `require_ratio_gate: true`
only when the declared contract uses these ratios. An explicit user choice may
instead use task-owned labels/flags or information-density closing; without any
declared maturity facts, maturity remains unknown.
Before launch, validate an actual file from the evaluator's canonical summary
writer with:
```bash
praxist resolve /path/to/task --result-summary /path/to/evaluation_summary.json
```
The check uses the runtime extractor and requires finite effort and coverage
ratios only when the task enables the ratio gate. Passing stage labels are not
required and cannot make missing ratio telemetry computable.
Canonical summaries must resolve to one completion decision under the
task-owned policy. Status vocabulary alone is not decisive: a fixed-budget
task may define reaching its configured cap as mature completion. The summary
must distinguish that case from an early stop through its achieved protocol,
effort/coverage, and completion fields. Task initialization tests both outcomes
through the real summary writer so the task policy, Frontier, Gems, and
prompt-facing views interpret the same result consistently.
Tasks with staged or full-coverage evaluation should define task-owned Gems
maturity through `selection_policy: mature_evidence_top_k` and
`min_mature_eval_units`. They may map task-owned stage labels to cumulative
evaluation-unit thresholds with `evidence_stage_min_units`; Praxist configuration
always calls these counts units, regardless of any task-local evaluator term.
Do not copy one task's threshold or stage labels into another task; derive both
from the complete protocol.
## Frontier Lanes And Incubator Evidence
`frontier/frontier_manifest.json` is the canonical machine-readable state for
the run's Frontier lanes and validation candidates. The manifest may contain
`lane_frontiers` when the task configures `evaluation.frontier_lanes`.
Operator-facing leaderboards are usually computed
from the SQLite finding store or from findings JSON fallback; a missing
standalone leaderboard file is not by itself evidence corruption.
Use frontier lanes when the task has staged validation or expensive full
evaluation:
```yaml
evaluation:
primary_metric: score
direction: maximize
frontier_lanes:
- name: confirmed
description: "Fully scored, promotable candidates."
k: 3
cumulative_cap: 10
axes:
- {name: score, direction: maximize}
include_lanes: [confirmed, performance]
require_metrics: [score]
parent_eligible: true
- name: incubator
description: "Lower-admission durable long-term library for task-authorized, protocol-passed, non-suspect Pareto/new-high candidates needing follow-up."
k: 8
cumulative_cap: 48
admit_new_high: true
axes:
- {name: score, direction: maximize}
# Put additional distinct metrics here when they should define
# Pareto/new-high retention and are emitted by every result mode the
# user-owned protocol authorizes for this lane.
optional_axes:
# Optional axes are secondary sort/display signals only; they do not
# by themselves define Pareto dominance.
- {name: secondary_tiebreak_metric, direction: maximize}
- {name: diagnostic_display_metric, direction: minimize}
include_lanes: [incubator, performance]
require_metrics: [score]
# This default excludes reduced modes. Adapt the structured filters when
# the user's protocol intent authorizes one of those modes as a parent.
require_falsey_metrics: [is_smoke_eval, partial, scout_only, validation_only, validation_only_result, late_after_generation_boundary, suspect_protocol, suspect_leakage]
parent_eligible: true
allow_non_promotable: true
allow_missing_tier: true
allow_risk_violating: true
- name: task_candidate
description: "Promising preliminary, aligned, or partial evidence retained for validation."
k: 5
cumulative_cap: 20
axes:
- {name: score, direction: maximize}
include_lanes: [task_candidate, candidate]
require_metrics: [score]
parent_eligible: false
allow_lower_tier: true
allow_non_promotable: true
allow_missing_tier: true
- name: diagnostic
description: "Controls, falsifiers, negative evidence, and process diagnostics."
k: 2
cumulative_cap: 10
axes:
- {name: score, direction: maximize}
include_lanes: [diagnostic, control, process, reference, negative_control]
parent_eligible: false
allow_lower_tier: true
allow_non_promotable: true
allow_missing_tier: true
```
Rename these lanes and metric requirements for the domain. The important
contract is not the specific names; it is that a lower-admission incubator
keeps task-authorized, protocol-passed Pareto/new-high variants available as
long-term parents, while modes marked non-parentable but still useful have a
declared candidate lane instead of being forced through the same gate as fully
promotable results.
`allow_lower_tier: true` retains lower-stage signals for revalidation; it does
not make that lane a durable parent source.
The lane name in an evaluator summary is a **source label**; each configured
lane is a durable target selected by Praxist. When confirmed and incubator should
both consider ordinary clean parent-authorized results, the evaluator should normally
emit one shared task-owned source label such as `performance`, and both targets
should list it in `include_lanes`. Do not map every such result directly to
`confirmed`: candidates outside confirmed top-k then cannot reach an incubator
that only accepts `incubator`/`performance`. At evaluator ingestion,
`frontier_lane`, `promotion_lane`, and `lane` are accepted source-label fields
in that precedence order; the first non-empty value is used. Committed entries
store the selected target in `frontier_lane` and `promoted_for_lane`, and
preserve a different submitted source label in `source_frontier_lane`.
Validation-candidate records use `submitted_frontier_lane`.
Task initialization must run a task-local reachability regression against real
summary construction. For each parent-eligible target, at least one
protocol-passed fixture from a parent-authorized mode must satisfy its source
and metric filters. With both confirmed and incubator present, test more than
`confirmed.k` parent-authorized candidates. When the task has a justified distinct incubator axis, include a
candidate that is outside primary top-k but non-dominated on that axis; it must remain incubator-eligible.
Do not invent a secondary metric for a genuinely single-metric task,
while fixtures from modes the task marks non-parentable, plus protocol-failed,
validation-only, late, and suspect fixtures, remain non-parentable. A temporarily empty incubator is valid when no
new Pareto point exists; an unreachable incubator is a harness defect.
Durable capacity is evidence-based rather than alias-based. Multiple finding or
variant names that reference the same exact immutable result artifact
(`source_result_path` and SHA-256) consume one durable lane slot. A different
path or hash is a different artifact, so independent replications remain
eligible. Semantic variant identity and lineage remain separate from this
capacity rule.
When non-code launch settings can alter a treatment, the evaluator summary
should additionally own a secret-free top-level `effective_config` object and
`effective_config_complete`. Include every task-owned argument, environment
override, protocol choice, and config-file value needed to reproduce the
treatment **after** the evaluator has applied defaults, aliases, parsing, and
type conversion. The resolved treatment is authoritative: omitting a setting
and explicitly supplying its resolved default must produce the same object and
digest. A genuinely different resolved value must produce a different digest.
Do not use an unfiltered process-environment snapshot as the contract. Praxist
carries a deterministic digest and the existing summary path into findings,
frontier, Gems, PI context, and reports; it does not copy the full object into
those derived views.
For derived work, a task-owned evaluator or existing launch helper should read
the selected parent summary and compare the child using this same resolved
schema before expensive execution. It may inherit allowlisted scientific values
when the task defines that behavior, but must not replay a complete parent
process environment or expose credentials.
For an exact replication, publish
`replication_of_effective_config_sha256` from the selected parent. Only a
completed result whose current complete digest matches that value supports the
exact-replication label. A result without these optional fields remains fully
compatible and follows the task's existing maturity and promotion rules.
## Evaluation Entrypoint
Task-owned evaluation may use any internal harness layout, but Praxist-facing
instructions should name a single public command. Prefer the structured
`task_entrypoints.evaluation.command` field; Praxist normalizes it into the
legacy-compatible `toolchain.eval_entrypoint` only in memory when needed:
```yaml
task_entrypoints:
evaluation:
command: evaluations/pareto_tiered/run.py
output_policy: compact stdout plus raw evidence under the run directory
```
Agents should call the evaluation entrypoint. Internal benchmark files under
`assets/harness/` are task implementation details and may change without
changing the Praxist task contract.
Keep task-owned evaluator, trainer, config, and harness paths relative to the
task root. The scheduler resolves a statically identifiable task-owned command
entrypoint against that root while preserving the task's configured working
directory, even when the operator starts or resumes Praxist elsewhere.
Absolute paths remain valid for external datasets, simulators, environments,
or services, but task
initialization must verify each one before launch. A new task is launch-ready
only after the declared interpreter runs the public evaluator successfully
from both the task root and a run-like working directory and resolves the same
task-owned paths in both cases.
Compact result summaries may be nested under `results/**/` and use
`summary.json`, `evaluation_summary.json`, `eval_summary.json`,
`tiered_eval_summary.json`, or `custom_*_tiered_eval_summary.json`;
`result_summary.json` is retained for compatibility. Put task-owned lane,
maturity, effort/coverage ratio, protocol, parent-use, and diagnostic metadata
in structured summary fields so materialization can preserve it in canonical
findings. Publish a stable top-level `variant_id` for each candidate (or an
explicit child-result ID when one evaluator emits several candidates) and reuse
that identity across evaluation stages; directory names remain fallback
provenance rather than candidate identity.
Praxist normalizes `task_entrypoints.evaluation.command` into the legacy-compatible
`toolchain.eval_entrypoint` when the latter is absent, then forwards that single
public evaluator to the default Claude SDK runtime. A direct Bash invocation is
registered through the existing protected-PID process-group launcher, so
generation drain and `active_evals` observe the same job. Keep the explicit
`protected_pids launch` form in task prompts whenever task-owned code launches
child work.
New tasks using central admission declare resource profiles and submit with a
stable semantic tag, profile, and work class. Profiles come from a timestamped
trace of the unchanged baseline rather than a synthetic CPU-versus-accelerator
comparison or a single teardown sample. Supply defaults, pressure rules,
process ownership, retries, and backend-specific handoff checks are defined in
[Central Experiment Scheduler](https://praxist.sapient.inc/en/docs/guides/central-resource-scheduler).
## Boundary Rules
- Praxist core discovers a task only after the CLI passes an explicit resolved
task-project path. The operator CLI resolves
`--task-path` > `TASK_PATH` > invocation directory.
- Task-local refs such as `task_role:*`, `task_audit:*`, and
`task_evaluation:*` are resolved inside the task project, not through the
global plugin catalog.
- The Praxist repo should not gain task-specific roles, audits, evaluations, or
harness code in `praxist/plugins/**`.
- A task project may ship its own generic plugins (e.g. a task-specific
`panel_topology`) under `/.praxist/plugins///`.
These are discovered with `source="task_project"` and take priority over a
same-named bundled plugin. They are only scanned when `--task-path` selects
the task; no implicit scan of arbitrary task directories happens.
## Task Tests
Task projects should include their own tests for the harness, fixtures,
evaluation profiles, and optional reference implementations. The Praxist repository
tests the task-project boundary and templates; it should not become the permanent
test suite for a private external task.
---
# Examples And Templates
Praxist ships two kinds of task-oriented assets. Their purposes are deliberately
different.
| Asset | Use it when | What it contains |
|---|---|---|
| Template | You are creating or testing a new task project | Replaceable scaffolding, placeholders, and deterministic smoke fixtures |
| Example | You want to inspect a finished integration | A complete task harness, evaluator, task-specific code, redistributable assets, evidence, and tests |
Templates under `templates/tasks/` are meant to be copied and adapted. Their
defaults illustrate contract shape; they are not universal scientific policy
and may intentionally omit datasets or production evaluators.
Examples under `examples/` are runnable reference projects. They retain their
own domain assumptions, metrics, resource profiles, and evidence boundaries.
Those choices demonstrate one project and must not become Praxist-wide defaults
or be copied uncritically into another field.
Both remain outside the `praxist` system package. Praxist core and generic
plugins contain no facts from either asset. Run lightweight checks in place,
but copy a template or example outside the Praxist checkout before starting a
research run so generated artifacts remain external to product source.
## Materialize A Complete Example
First-use setup creates writable copies under
`${PRAXIST_EXAMPLES_HOME:-~/PraxistExamples}`. Inspect or recreate them with:
```bash
praxist examples list
praxist examples install rocket_booster_recovery
praxist examples install rocket_booster_recovery_rust
```
Pass `--destination /absolute/path` to install one example elsewhere. Existing
destinations are preserved unless the operator explicitly chooses another
path.
## Choose A Starting Point
- Start with a [task template](https://praxist.sapient.inc/en/docs/reference/task-templates) when authoring a
new task contract.
- Study [Rocket Booster Recovery](https://praxist.sapient.inc/en/docs/examples/rocket-booster-recovery) for a
Python/JAX integration or [Rocket Booster Recovery
(Rust)](https://praxist.sapient.inc/en/docs/examples/rocket-booster-recovery-rust) for an offline native
Rust integration of the same research problem.
- Read [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects) for the canonical task contract shared
by both.
---
# Rocket Booster Recovery
Rocket Booster Recovery is a complete classical-control example with a frozen
six-degree-of-freedom plant, deterministic data banks, a first-contact landing
evaluator, task-local research roles, and preserved baseline evidence. It shows
how a real research project and its Praxist task harness fit together without
moving domain logic into the framework.
The source checkout contains the example at:
```text
examples/rocket_booster_recovery/
```
A Praxist wheel exposes the same tree as the package resource
`praxist/resources/examples/rocket_booster_recovery/`.
## What It Demonstrates
- a complete task harness rather than a scaffold;
- one public evaluator backed by a frozen simulation boundary;
- task-owned metrics, maturity rules, frontier lanes, roles, and audit rules;
- measured evidence kept separate from unmeasured platform claims;
- small frozen data banks and manifests protected by checksums;
- lightweight contract tests that do not launch a research run.
The controller design, physical constraints, metrics, and hardware profiles are
specific to this example. They are not Praxist defaults. See
[Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates) before adapting
any part of it.
## Inspect And Test
The source and package-resource copies are inspection-only. Materialize a
writable project before creating an environment or running any code:
```bash
praxist examples install rocket_booster_recovery
cd ~/PraxistExamples/rocket_booster_recovery
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
./scripts/run_tests.sh
```
The project README owns the scientific protocol, environment details, baseline
measurements, and evaluator commands.
## Start A Research Run
The first `praxist setup` after pip installation creates a writable copy at
`~/PraxistExamples/rocket_booster_recovery` and prints the absolute path.
Recreate it explicitly or choose another destination with:
```bash
praxist examples install rocket_booster_recovery
praxist examples install rocket_booster_recovery \
--destination /path/to/rocket_booster_recovery
```
An existing destination is preserved without replacement. Prepare that working
copy's environment, then choose the task harness that matches the host:
| Harness | Intended host | Evidence status |
|---|---|---|
| `task_GPU_server` | A compatible accelerator server | Includes the project's measured server baseline |
| `task_PC` | A workstation or laptop | Requires a baseline measured on that machine before research |
```bash
cd ~/PraxistExamples/rocket_booster_recovery
praxist resolve "$PWD/task_GPU_server" --run-dir "$(mktemp -d)"
praxist start --task-path "$PWD/task_GPU_server" --daemonize --json
```
The bundled server evidence must not be relabeled as evidence from another
machine. The project README owns hardware details and scientific protocol.
---
# Rocket Booster Recovery (Rust)
This complete example implements the Rocket Booster Recovery research problem
as a native Rust project. It is independent of the Python/JAX example, keeps
its Cargo dependencies vendored for offline builds, and includes three Praxist
task profiles:
| Profile | Intended host |
|---|---|
| `task_GPU_server` | High-concurrency Linux server; the evaluator remains CPU-only |
| `task_linux` | Portable x86_64 or aarch64 Linux host |
| `task_macos` | Apple Silicon macOS host with platform-specific baseline qualification |
The packaged copy is read-only. Install a writable project before building,
evaluating, or launching research:
```bash
praxist examples install rocket_booster_recovery_rust
cd ~/PraxistExamples/rocket_booster_recovery_rust
```
The project requires Rust 1.85 or newer. Validate the selected profile from the
writable copy:
```bash
cargo test --release --workspace --locked --offline
praxist resolve "$PWD/task_linux" --run-dir "$(mktemp -d)"
```
Then launch through an agent skill or the direct CLI using that task directory:
```bash
praxist start --task-path "$PWD/task_linux" --daemonize --json
```
The example's [source README](https://github.com/sapientinc/praxist/tree/main/examples/rocket_booster_recovery_rust#readme)
is the sole owner of its scientific protocol, baseline evidence, build audit,
platform constraints, and task-profile details.
---
# Direct CLI Operations
This guide is for operators who want to control Praxist directly from a shell,
without asking an agent to perform the lifecycle action. For guided operation,
describe the desired action in Codex or Claude Code; it will use the
`praxist-control` skill when appropriate.
The `praxist` CLI is the canonical shell interface. The Python module entrypoint
is a low-level compatibility surface.
```mermaid
flowchart LR
VALIDATE(["Validate
doctor + resolve"])
RUNNING["Running
start --daemonize"]
OBSERVE(["Observe
status / monitor"])
BOUNDARY[["Lifecycle boundary
stop / resume / complete"]]
VALIDATE --> RUNNING --> OBSERVE
OBSERVE -.-> RUNNING
RUNNING --> BOUNDARY
BOUNDARY -.-> RUNNING
class VALIDATE,OBSERVE interface
class RUNNING system
class BOUNDARY artifact
```
## Command Map
| Goal | Direct command |
|---|---|
| Reopen first-use runtime setup | `praxist setup --interactive` |
| Hand a project to guided takeover | `praxist --takeover --task-path ` |
| Check host and runtime readiness | `praxist doctor --task-path ` |
| Validate a task without starting | `praxist resolve ` |
| Start a detached run | `praxist start --task-path --daemonize --json` |
| List or inspect runs | `praxist status --json` |
| Inspect one run | `praxist status --run-id --json` |
| Open the read-only TUI | `praxist --monitor --run-id ` |
| Stop one run | `praxist stop --grace 300 --json` |
| Resume a clean interrupted run | `praxist resume --daemonize --json` |
Use `praxist --help` for the live argument contract. The generated
[CLI Reference](https://praxist.sapient.inc/en/docs/reference/cli) is derived from that parser.
The first command reopens setup. The second is a separate, post-manual project
handoff described in the [Quickstart](https://praxist.sapient.inc/en/docs/getting-started/quickstart). Neither
replaces the task validation or lifecycle commands below.
## Select the Task and Configuration
An explicit `--task-path` wins over `TASK_PATH`, which wins over the invocation
directory. Prefer the explicit form in scripts:
```bash
praxist resolve /absolute/path/to/task
praxist start --task-path /absolute/path/to/task --daemonize --json
```
Lifecycle commands load `${XDG_CONFIG_HOME:-$HOME/.config}/praxist/env` by
default. When using another configuration file, pass it to every
gate and lifecycle action:
```bash
praxist doctor --task-path /absolute/path/to/task --config-file /path/to/env
praxist resolve /absolute/path/to/task --config-file /path/to/env
praxist start \
--task-path /absolute/path/to/task \
--config-file /path/to/env \
--daemonize \
--json
```
Use the same configuration for validation and startup. Explicit
process environment values have the highest credential precedence. See
[Credentials](https://praxist.sapient.inc/en/docs/guides/credentials) for secret handling and
[Agent Runtimes](https://praxist.sapient.inc/en/docs/guides/agent-runtimes) for runtime selection.
For a source checkout, run the same CLI from the repository environment:
```bash
cd /path/to/Praxist
uv run praxist start \
--task-path /absolute/path/to/task \
--daemonize \
--json
```
??? note "Shared state and host identity"
Registry actions record a best-effort host identity so a shared state
directory cannot make one machine stop another machine's run. A minimal or
rebuilt container without a stable machine ID should set one persistent,
host-local `PRAXIST_HOST_ID`. Never reuse that value across distinct hosts.
## Validate and Start
### Configured API Provider
Run the inexpensive gates before launching:
```bash
praxist doctor --task-path /absolute/path/to/task
praxist resolve /absolute/path/to/task
praxist start \
--task-path /absolute/path/to/task \
--daemonize \
--json
```
Preserve explicit `--runtime`, `--model-provider`, `--model`, `--cohort`,
`--generations`, and `--strategy` choices when required. `--daemonize` lets the
run survive the launching shell or Codex session. `--json` makes run ID, PID,
run directory, log path, and monitor handoff machine-readable.
### Codex-native Mode
Use the route-aware doctor and pass the same mode to start:
```bash
praxist doctor --codex-native --task-path /absolute/path/to/task
praxist resolve \
/absolute/path/to/task \
--codex-native \
--runtime agent_runtime:codex_sdk \
--model-provider model_provider:openai_compatible \
--model gpt-5.6-luna
praxist start \
--codex-native \
--task-path /absolute/path/to/task \
--agent-system codex_sdk \
--runtime agent_runtime:codex_sdk \
--model-provider model_provider:openai_compatible \
--model gpt-5.6-luna \
--daemonize \
--json
```
The saved-login route is isolated from configured relay API providers and API-key
endpoints. Use `praxist-takeover-codex` for the guided equivalent.
## Observe a Run
### Status
```bash
praxist status --json
praxist status --run-id --json
```
The targeted form is preferable in automation. It avoids mixing unrelated runs
when multiple task projects share a host.
### Foreground Monitor
```bash
praxist --monitor --run-id
```
The fullscreen TUI is read-only and independent of the research process. It
shows run state, peer health/activity, recent log context, and hardware
warnings. Visual redraw is decoupled from bounded artifact and hardware
sampling, so display responsiveness does not multiply research-side probes.
`Ctrl-C` exits only the monitor; the detached run continues. Use
`praxist --monitor --plain` for a non-interactive terminal or append-friendly
transcript.
For peer rows, the live monitor reads each peer's bounded
`recent_result_artifacts` summary instead of recursively reconciling the complete
result tree in the long-running display. Use `praxist --monitor --once`,
`praxist status`, or the diagnostic workflow when you need a complete artifact
reconciliation view.
`--interval` controls the display interval. Exact defaults and limits are owned
by the generated [CLI Reference](https://praxist.sapient.inc/en/docs/reference/cli).
### Interpret Operational State
| View | Authority |
|---|---|
| Result and finding summaries | Measured task evidence |
| `frontier/frontier_manifest.json` | Canonical lane and promotion state |
| Committed `gems/gems_state.json` | Canonical Gems state |
| `gen_N/generation_boundary.json` | Canonical generation boundary |
| Leaderboards, Principal Investigator (PI) evidence packs, rendered prompts, reports | Derived views or audit snapshots |
| TUI and scheduler status | Live operational telemetry, not scientific evidence |
Count completed generations from the contiguous committed boundary markers.
If live status or `run_summary.json` reports a larger value than those markers,
report the mismatch and classify the extra generation as pending boundary work;
do not treat frontier entries or `generation_results.json` alone as a commit.
When central scheduling is enabled,
`/resource_scheduler/status.json` distinguishes queued, running, blocked,
completed, failed, and rejected work. Read lifecycle `running` separately from
`running_activity.by_resource_phase`; a live wrapper is not proof of active
accelerator compute, and an observation marked `unknown` is not proof of a
stall. Resource telemetry helps explain throughput but cannot promote a result.
## Stop a Run
Target one verified run whenever possible:
```bash
praxist status --run-id --json
praxist stop --grace 300 --json
```
After stop returns, poll status until the selected process disappears. Use
`praxist stop --all` only when every registered run owned by the environment is
intentionally being stopped. Do not use broad `pkill` patterns.
The foreground monitor is independent, so stopping a run does not need to find
or kill a monitor process.
## Resume a Run
For a run stopped at a clean, recognized boundary:
```bash
praxist status --run-id --json
praxist resume --daemonize --json
```
Never resume a verified live controller. Praxist preserves the original API
provider, agent runtime, model, frontier strategy, and task identity; resume
rejects changes to these canonical values. An unchanged task checkout may move
to a new absolute path; Praxist validates its persisted manifest and effective
descriptor. Task identity comes from the persisted task contract.
Interrupted final boundaries require more care. Common cases include:
- an unfinished final generation after a complete PI agenda;
- a finished cohort whose PI/Chair boundary did not finish;
- a committed Gems reset followed by a partial next generation;
- a pending or incomplete Gems reset transaction.
The operator agent should prepare the run directory before calling the resume command for
these irregular cases. It should inspect the Praxist resume plan, back up the
run before any manual crop, preserve complete generation evidence and committed
Gems state, and prefer Praxist's internally recoverable boundary path. When a
partial boundary is not recognized, do not hand the partial state directly to `praxist resume`;
use a documented repair path or crop to a named clean boundary only with
operator approval.
Use `praxist-control` with the request "resume the latest run" for this preparation workflow. The
skill understands interrupted PI and Gems boundaries and avoids destructive
guessing.
## Agent-Assisted Operation
Codex or Claude Code is the recommended interface when an action requires path selection,
artifact interpretation, irregular resume preparation, or a concise progress
report. Ask in natural language, for example:
```text
Report current research progress and list the strongest variant in every
completed generation with its task-defined performance metrics.
```
The agent will use `praxist-control` as needed. If invoked without an operation,
the control skill asks for `start`, `stop`, `resume`, `status`, `monitor`, or
`detect-active-runs` instead of guessing.
For launch, the agent must know the exact task project before launching. It should
confirm the exact path. It should not infer a task from a broad filesystem scan.
During a status request it reports generation progress, incubator or leaderboard
performance, CPU/memory/process/accelerator load, generated report paths, and at
most two score curves through `terminal-line-plot`. It must not stop, resume,
crop, rerender, or edit files during a status request. Canonical state remains
authoritative; derived reports remain audit snapshots.
### Guided Diagnostics and Reports
Use `praxist-diagnostic` when the question is why a run is unhealthy or weak,
not merely what state it is in. The default diagnostic is analysis-only and may
write a report under task `docs/`; it must not edit task code or active run
artifacts. It can audit diversity HHI (Herfindahl-Hirschman Index), artifact
consistency, resource/runtime
friction, a performance ceiling, and the strongest variants. For persistent
weakness it can build a chronological agent behavior analysis report. Manual
A/B/C run reports put strongest results first, strong-variant lineage second,
and run health third. These reports are derived views: canonical state remains
authoritative and report snapshots remain audit snapshots.
## Output Locations
`praxist start --json` returns the selected run and launcher-log paths;
`praxist status --json` resolves registered runs without relying on the current
directory. [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects#experiments-directory) owns output
placement, run contents, and task-runtime path rules.
---
# Run Reports
Run reports are human-readable, derived views of canonical Praxist evidence.
They explain progress without becoming another result, frontier, or baseline
owner.
```mermaid
flowchart LR
EVIDENCE[("Canonical run evidence")]
REPORT["Markdown + PDF report"]
READER(["Researcher / diagnostic agent"])
EVIDENCE --> REPORT --> READER
class EVIDENCE artifact
class REPORT system
class READER actor
```
## Automatic Reports
Praxist writes reports under `/docs/praxist_reports/` when:
- the first credible frontier result beats a known baseline on the same metric;
- the completed-generation count reaches a multiple of three (3, 6, 9, and so
on); or
- the run reaches a terminal state.
The generated subtree is excluded from task identity, so report refreshes do
not change the task manifest or block a compatible resume.
## Report Structure
Every report follows the same order:
1. **Strongest variants or Pareto front** presents task-declared metrics,
credibility, and concise mechanism summaries. With no clean frontier entry,
dimension winners from lower-authority evidence may appear only as clearly
labeled signals.
2. **Strong-variant lineage** follows only those variants through parents,
generations, source results, and inherited ideas.
3. **Run health** summarizes artifact consistency, operational friction,
diagnostic coverage, and caveats.
The PDF companion uses the same facts and adds charts only when task-declared
numeric directions support them. A metric with unknown direction may be shown
as context, but it cannot select a winner, trigger a baseline claim, or drive a
directional chart.
## Generate a Report
`tool_server:run_report` exposes `generate_run_report` inside an agent session.
The `praxist-diagnostic` skill can generate the same report during an explicit
run-health analysis.
Canonical truth remains in result and finding summaries,
`frontier/frontier_manifest.json`, committed `gems/gems_state.json`, and
`gen_N/generation_boundary.json`. Reports never feed promotion or generation
close decisions.
---
# Cost Estimation
Praxist can spend tokens, task-selected compute time, wall-clock time, tool quota,
and external API quota. Cost estimates are advisory, not a replacement for
BudgetPolicy.
## Inputs To Estimate
Estimate cost from:
- cohort size;
- number of generations;
- model profile per stage;
- expected peer session count;
- expected Principal Investigator (PI) and Chair planning calls;
- prompt size and cache stability;
- tool calls;
- evaluation runtime;
- platform/backend capacity and queue behavior, including accelerator capacity
only when the task actually uses one.
## Prompt Cache Readiness
PromptLayout V1 keeps frozen and dynamic prompt blocks separate. Stable frozen
prefixes improve the chance that caches managed by the agent runtime or API
provider are useful.
Agent runtime cache behavior is currently treated as runtime-managed. Praxist records
layout hashes and cache provenance rather than injecting raw cache directives
where the runtime does not expose them.
## Interpreting Usage
Exact token or cache usage depends on what the agent runtime/API provider returns. Missing
metering should be recorded as unknown rather than zero.
Use run artifacts, budget ledgers, and API provider invoices together when analyzing
cost after a dogfood run.
For runtimes that report cache usage, `input_tokens` is the inclusive logical
input total and `cached_input_tokens` is the cache-read subset. Calculate:
```text
uncached_input_tokens = input_tokens - cached_input_tokens - cache_creation_input_tokens
cache_hit_ratio = cached_input_tokens / input_tokens
sessions_per_peer_generation = peer session count / peer-generation count
```
Treat an unreported cache-creation value as zero. If the reported components do
not fit inside inclusive input, keep the raw values and mark them inconsistent
instead of forcing the equation to balance.
Claude SDK telemetry additionally preserves cache-creation input separately.
Its API provider's native `input_tokens` value is uncached input, so the adapter
normalizes inclusive input as uncached + cache read + cache creation. Historical
or third-party records whose cached input exceeds their declared total are kept
unchanged and marked `telemetry_inconsistent`; Praxist does not clamp them or
publish a misleading cache-hit ratio.
Treat these as separate signals. Prompt caching can reduce billed compute while
logical input remains high; reducing unnecessary fresh sessions reduces both
repeated tool work and logical input. The read-only diagnostic inventory derives
these values from canonical `generation_results.json` rows and the run summary;
it does not create another usage ledger.
Session-reuse and tool-output mechanisms are defined in
[Cost Optimization](https://praxist.sapient.inc/en/docs/guides/cost-optimization); this page only defines how to
estimate and interpret cost.
---
# Architecture
Praxist is a task-agnostic autonomous-research control plane:
```text
stable core contracts + generic plugins + explicit external task projects
```
```mermaid
flowchart LR
TASK[["Task project
domain truth + evaluator"]]
ENTRY(["CLI + resolver
frozen run configuration"])
ENGINE["Praxist
core contracts + generic plugins"]
RUN[("Task-local run
canonical evidence + derived views")]
TASK --> ENTRY --> ENGINE --> RUN
class TASK task
class ENTRY interface
class ENGINE system
class RUN artifact
```
## Ownership Boundary
| Owner | Responsibilities | Must not own |
|---|---|---|
| Praxist core | Stable protocols, resolution, canonical storage, replay, credentials, budgets, and extension interfaces | Task facts, provider wire objects, or workflow implementation details |
| Generic plugins | Replaceable runtimes, API providers, workflow stages, tools, graph maintenance, topology, and budget policy | Facts or prompts usable only by one task |
| Task project | Objective, baseline, evaluator, metrics, evidence policy, roles, prompts, assets, and scientific constraints | Praxist source or another task's state |
| Run directory | Frozen configuration, artifacts, evidence, lifecycle state, and replay records for one run | Mutable task configuration as current truth |
The only system package is `praxist`. A component is generic only when two
unrelated task projects can use it without importing one task's private facts.
The complete task contract is defined in [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects);
templates and complete examples are classified in
[Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates).
Run artifacts never belong in the Praxist checkout. [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects#experiments-directory)
owns output-path selection and the task-local run boundary.
## Core and Plugin Boundary
`praxist.core` supplies protocols and interfaces for task resolution, runtime
requests and events, model profiles, tools, artifacts, findings, trajectory,
budget, credentials, workflow stages, topology, and replay. Selectable behavior
belongs behind one of those interfaces.
Executable generic plugins live under `praxist/plugins/**` or an explicitly
selected plugin root. Their manifests identify compatibility, entrypoints, and
source-hashed code or assets. Task-local generic plugins can live under
`/.praxist/plugins/` and are visible only when that task is selected.
[Generic Plugins](https://praxist.sapient.inc/en/docs/guides/plugins) owns manifest and testing rules.
[Configuration Discipline](https://praxist.sapient.inc/en/docs/concepts/config_discipline) owns configuration ingress and
precedence. [Panel Topology Prompts](https://praxist.sapient.inc/en/docs/concepts/panel_topology_prompts) owns the
prompt-asset loader contract.
## Runtime and Workflow Boundaries
Agent runtimes consume normalized requests and return normalized events/results.
API provider plugins describe API shape, models, credentials, and route
capabilities; core calls neither a concrete SDK nor provider directly.
[Agent Runtimes](https://praxist.sapient.inc/en/docs/guides/agent-runtimes) and
[API Providers](https://praxist.sapient.inc/en/docs/guides/model-providers) own those contracts.
`workflow_stage:research_loop` owns the executable research loop. Its topology
sidecar records each generation's cohort, while the module API exposes read
views and records structured external requests. The bundled executor does not
apply queued topology mutations. The detailed boundaries live in
[Workflow Stages](https://praxist.sapient.inc/en/docs/guides/workflow-stages) and
[Research Topology Audit API](https://praxist.sapient.inc/en/docs/guides/research-topology-and-module-api).
[Central Experiment Scheduler](https://praxist.sapient.inc/en/docs/guides/central-resource-scheduler) owns
experiment admission and resource allocation inside the research-loop plugin.
## State and Replay
Run artifacts have one of five roles:
- **`canonical_state`** is current machine-trusted state, including measured
result/finding evidence, committed frontier and Gems state, and generation
boundaries.
- **`validation_signal`** is useful non-durable evidence retained for repair,
diagnosis, or follow-up. Task policy decides whether later canonical evidence
may promote it.
- **`derived_view`** is a regenerable bounded view, such as a leaderboard or
run report.
- **`audit_snapshot`** records what a stage or agent saw, including rendered
prompts and evidence packs.
- **`partial_output`** is interrupted or rejected output and is ignored by
normal runtime readers.
The rule is fewer fact owners, not fewer signals. Planning, promotion, reset,
resume, and close read canonical state plus explicitly eligible signals.
Derived views and audit snapshots explain decisions but never override their
sources.
`gen_N/generation_boundary.json` is the commit acknowledgement for a
generation. Results without a contiguous boundary remain pending work. A final
evidence cutoff prevents late files from rewriting a closed generation; they
remain visible as follow-up signals. Auto-materialized findings are idempotent,
rebuildable projections of result summaries rather than a second result owner.
The [Research Loop](https://praxist.sapient.inc/en/docs/guides/research-loop-variant-generation-flow) owns the
write order and inheritance sequence.
## Finding Graph
The finding graph is advisory research context. Graph maintainers write edges,
health, and compact guidance; query tools expose those views to peers and
panels. The graph cannot rewrite raw findings, frontier membership, or task
ranking. Alternative graph implementations belong behind the graph-maintainer
plugin contract.
## Result Preservation
Praxist favors survivable long-running research:
```text
capture first, label uncertainty, continue when safe
```
Useful output is retained unless continuation would corrupt the fact chain,
expose secrets, damage existing results, or exceed an approved resource
envelope. Weak provenance is labeled; missing usage is `usage_unknown`, never
silently zero.
## Verification Boundary
`AGENTS.md` is the contributor contract and the generated reference is the
enumerated API surface. Default tests remain offline: they do not require real
model keys, network, GPUs, external task repositories, or long research runs.
---
# Research Loop
Praxist turns one external task project into successive generations of
candidate implementations and measured evidence. Core does not compile a task
prompt into code: peers inspect the baseline, implement hypotheses, run the
task-owned evaluator, and publish evidence. Praxist coordinates that work and
commits what later generations may inherit.
Principal Investigator (PI) agents propose next-generation work; multi-PI
topologies add a Chair that consolidates those proposals. The
[Glossary](https://praxist.sapient.inc/en/docs/about/glossary) gives compact definitions, while
[Panel Topology Prompts](https://praxist.sapient.inc/en/docs/concepts/panel_topology_prompts) owns the panel
contract.
```mermaid
flowchart LR
PLAN[["Task contract +
committed agenda"]]
RESEARCH["Parallel peers +
task-owned experiments"]
EVIDENCE[("Results + findings +
retention lanes")]
PANEL["PI agents / Chair +
next agenda"]
PLAN --> RESEARCH --> EVIDENCE --> PANEL
PANEL -.-> PLAN
class PLAN task
class RESEARCH,PANEL system
class EVIDENCE artifact
```
## 1. Resolve and Freeze the Run
Startup resolves the selected task, plugins, prompts, baseline references, API
provider, agent runtime, and initial durable state. It writes a run-local frozen
configuration before cohort execution. The task schema and precedence rules are
defined in [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects); the research loop consumes that
contract without adding domain meaning.
## 2. Build the Generation Context
`GenerationLoop` combines:
- the task prompt, peer role, and allowed work surface;
- the committed agenda and optional Deep Innovation Gate (DIG) or
Quality-Diversity (QD) allocation;
- compact frontier, incubator, Gems, graph, and negative-evidence views;
- peer-local research memory; and
- the task-owned evaluator and protocol contract.
Generation zero starts from the baseline and initial task context. DIG may run
once before its cohort when enabled; QD is independently selectable. Later
generations use the preceding committed PI/Chair agenda. DIG is normally off,
while QD can allocate candidate contracts through the existing synthesis path.
[Deep Innovation Gate](https://praxist.sapient.inc/en/docs/guides/deep-innovation-gate) and
[Quality-Diversity Allocation](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator) own those mechanisms.
The workflow records `gen_N/research_topology.json` so worker identities,
declared inputs/outputs, and visibility policy remain auditable without changing
peer semantics.
## 3. Execute Peer Work
Each peer receives a rendered prompt and a normalized runtime request. A typical
peer:
1. reads its task, role, agenda, and inherited evidence;
2. states a mechanism hypothesis and intended evidence stage;
3. creates an independent variant under `variants/`;
4. changes only permitted files;
5. submits evaluation through the task's public evaluator path;
6. writes structured output under `results/`; and
7. publishes a finding with caveats and follow-up.
The selected resource policy controls experiment admission. It does not choose
the hypothesis or change scientific validity. See
[Central Experiment Scheduler](https://praxist.sapient.inc/en/docs/guides/central-resource-scheduler).
## 4. Materialize Evidence
A **result summary** answers what the evaluator measured. A **finding** records
how later research may interpret and use that result. Praxist recursively
discovers recognized summaries and idempotently materializes their usable
metadata into canonical findings.
Findings preserve actual evidence stage, maturity ratios, metrics, mechanism,
caveats, lane intent, parent eligibility, and follow-up. A repeated finding ID
can refresh changed non-empty fields without erasing useful older fields omitted
from the update.
Committed lane membership lives in
`frontier/frontier_manifest.json`. A leaderboard is only a derived view.
Task-owned lane, maturity, and Gems policy is defined in
[Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects#frontier-lanes-and-incubator-evidence).
Immature scout or partial output
remains a validation signal, not incubator content.
Artifact-role definitions are owned by
[Architecture](https://praxist.sapient.inc/en/docs/concepts/architecture#state-and-replay). The loop plans
from canonical state, keeps useful validation signals visible, and never treats
old prompts, reports, or evidence packs as current truth.
## 5. Close and Commit the Generation
After admitted work drains, the boundary performs one ordered commit:
1. ingest peer findings and result summaries;
2. run a final idempotent evidence refresh;
3. update canonical finding, graph, frontier, incubator, and optional Gems state;
4. refresh peer memory and negative-evidence summaries;
5. build the PI evidence view and synthesize the next agenda; and
6. write `gen_N/generation_boundary.json`.
The marker is the completion fact. Results or frontier files without a
contiguous marker are pending boundary work for resume to finish.
A recorded evidence cutoff makes retries deterministic. Results published after
that cutoff remain visible as late validation signals, but cannot enter the
closed generation through retry timing. Atomic files visible before the cutoff
remain eligible even if ingestion observes them during reconciliation.
[Research-Loop Flexibility Controls](https://praxist.sapient.inc/en/docs/guides/research-loop-flexibility-controls)
owns close eligibility, mature quorum, drain, and bounded-liveness behavior.
## 6. Synthesize and Inherit
PI/Chair synthesis reads committed evidence, prior agenda, task constraints,
memory, and diversity diagnostics. It writes the next agenda under
`agendas/research_agenda_gen.yaml`.
An agenda assigns planned work; it is not measured evidence. Evaluator results
and committed retention state continue to own scores, maturity, protocol status,
and parent eligibility. A failed or uncommitted agenda cannot drive another
cohort.
A later peer may restart from the baseline, inherit a durable candidate, repair
a credible signal, ablate a strong result, combine compatible mechanisms, or
investigate an anti-mainline direction. Praxist supplies evidence and
constraints, not a fixed code-generation template.
## 7. Audit the Flow
| Question | Canonical location |
|---|---|
| What was implemented? | `variants//` |
| What was measured? | `results//` |
| What evidence was published? | `findings/`, `shared_findings/`, or the canonical finding store |
| What was durably retained? | `frontier/` and `gems/` |
| What did the panel plan next? | `agendas/` |
| Did the generation commit? | `gen_N/generation_boundary.json` |
A variant present only under `variants/` or `results/` may not influence
later planning if no usable finding can be published or materialized.
---
# Research-Loop Flexibility Controls
These controls preserve useful signals while making maturity, retention, and
generation close follow the task owner's declared protocol. Exact task fields
and the recommended combined profile are defined in
[Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects#deep-innovation-gate-quality-diversity-and-gems-defaults).
## Mature Evidence
When a task enables ratio-based maturity, its canonical evaluator summary emits:
- `effort_ratio`: actual evaluation effort divided by the task's mature
reference effort;
- `coverage_ratio`: completed required evaluation units divided by all
required units.
Praxist copies these normalized facts into auto-materialized findings. A
standalone finding without a canonical summary reference must carry them
itself. With `require_ratio_gate: true`, missing or non-finite ratios remain
unknown; stage names cannot fill the gap. A task that deliberately uses labels
or completion flags instead may leave ratio gating disabled and define those
semantics explicitly.
Initialization validates one real file from the evaluator's production summary
writer with `praxist resolve --result-summary`. During a run, a durable-looking
result missing required ratios remains visible as a validation signal and
produces one bounded warning; it is not promoted or counted for mature close.
Gems, frontier lanes, reports, and close all consume the same maturity decision.
Praxist counts generic evaluation units and never gives a domain-specific stage
name global meaning.
## Generation Close
`synthesis_trigger.mature_quorum_fraction` controls normal close when a task
distinguishes close-grade evidence:
| State | Behavior |
|---|---|
| Positive quorum reached | Freeze new work, drain active work, synchronize evidence, and close normally. |
| Positive quorum missing at assessment | Fence ordinary admission while deadline-safe mature top-ups remain eligible. |
| Quorum `0.0` | Information density may close normally; valid only when this is the task owner's intended protocol. |
| Safety bound or fully drained cohort | Preserve liveness and record insufficient maturity where applicable. |
The scheduler's mature supply target is advisory and cannot replace this gate.
When close begins, `CLOSING_SIGNAL` blocks every new experiment while already
running protected work finishes. A bounded drain grace lets agent sessions
publish final evidence before `STOP_SIGNAL`; it never kills an active protected
evaluator.
Runtime-owned stdout files are not completion facts because a successful command
may emit no text. Structured runtime completion and exit status own that state.
Task-owned progress files remain valid only when their readiness contract is
explicit.
The committed boundary records close reason, maturity outcome, and optional
peer-mix telemetry. `orchestrator_status.json` exposes that compact state to
status, monitor, and diagnostics.
## Durable Incubator
An incubator lane is a lower-admission, long-term library, not a stricter winner
lane. It can retain protocol-authorized, protocol-passed, non-suspect candidates
that establish a task-defined Pareto point or new high even when the confirmed
lane is full.
Confirmed and mature incubator lanes may be parent-eligible. Preliminary,
diagnostic, suspect, protocol-failed, and other task-declared non-parent modes
remain in validation lanes for follow-up. Optional display/tiebreak axes do not
silently become Pareto dimensions. The complete lane schema, source-routing
rules, and reachability test are owned by
[Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects#frontier-lanes-and-incubator-evidence).
Promotion rejection summaries remain inside
`frontier/frontier_manifest.json`; Praxist does not create a second runtime
fact file.
## Constructive Peer Mix
When enabled, Praxist estimates constructive solution work versus
diagnostic/control work at each committed boundary. The next generation sees
the result as advisory feedback, not a quota or execution gate. Disabling the
feature stops both calculation and prompt injection, including historical
telemetry on resume.
This control is independent of the Deep Innovation Gate (DIG) initial
innovation-slot policy and Quality-Diversity (QD) allocation. Those switches
are described in
[Deep Innovation Gate](https://praxist.sapient.inc/en/docs/guides/deep-innovation-gate) and
[Quality-Diversity Allocation](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator).
## Launch Freeze
The launch guard cooperates with the central scheduler to freeze queued and new
submissions before `CLOSING_SIGNAL`. It covers training, evaluation, scripts,
shell launchers, and background processes while still allowing result reading,
finding publication, and memory updates. Existing protected jobs drain
naturally.
Tasks may disable the guard only when they explicitly own an equivalent
close-safe boundary. Disabling it does not waive timing feasibility or mature
close requirements. Heavy-work and close-grade runtime estimates remain
separate because a task may permit long optional evaluation while authorizing a
shorter protocol for normal close.
Central submission, semantic retry, queue ownership, and process-group behavior
are defined only in [Central Experiment Scheduler](https://praxist.sapient.inc/en/docs/guides/central-resource-scheduler).
---
# Deep Innovation Gate
The Deep Innovation Gate (DIG) is a deep-reasoning innovation process that
compares mechanism-level alternatives before implementation.
DIG is not an experiment loop. It does not train, evaluate, write variants, or
encode task metrics. It uses the selected Praxist runtime/API provider and the
task's existing prompt, baseline, and file boundaries.
## Generation Scope and Flow
The recommended profile runs DIG only before absolute generation zero. Later
generations use committed agendas from Principal Investigator (PI) agents and,
in multi-PI mode, a Chair; Gems resets do not reactivate DIG.
For an enabled generation:
1. build the normal peer context;
2. map the baseline and generate/critique candidate mechanisms with read-only
planner tools;
3. validate one selected contract;
4. add that contract as a dynamic prompt block; and
5. launch the ordinary implementation peer.
The selected contract identifies the variant, mechanism, intervention surface,
rejected alternatives, planned files/changes, expected metric signature,
ablation hooks, and fail-fast checks. A material implementation deviation
requires an auditable `contract_amendment.yaml`.
## Artifacts
```text
gen_/peers//dig/
baseline_mechanism_map.yaml
candidate_pool.yaml
candidate_reviews.yaml
qd_selection.yaml
selected_contract.yaml
dig_summary.md
```
These are design/audit artifacts, never empirical findings. Result findings may
reference their metadata after evaluation. With `generation_scope:
initial_only`, their absence after generation zero is expected.
## Retry and Fallback
Malformed planner output or an invalid candidate/contract retries within the
configured attempt and total-time bounds. Valid phase checkpoints can be reused
only when prompt and artifact fingerprints still match.
After all attempts fail, Praxist writes `dig_failure_summary.json`. The default
then starts the ordinary implementation path, preserving liveness and an
audit trail. Strict tasks may disable fallback, accepting that one planning
failure can suppress a peer.
## Control Surface
Task initialization can enable or disable DIG independently, limit it to the
initial generation, bound planning time and candidate breadth, and choose
whether planner failure falls back to direct implementation. These controls
affect pre-code reasoning only; they do not change task evaluation or evidence.
## Validation
The gate requires enough mechanism and intervention diversity, critiques for
every candidate, at least one falsifying or diagnostic alternative, and a
complete selected contract. Evaluator, data split, and metric-calculation
changes are forbidden by default. Lane fit and duplicate checks apply only when
their corresponding task policies are active.
This validation predicts whether a plan is coherent. It never claims measured
performance or makes the contract a parent.
## Relationship to Quality-Diversity and Gems
DIG and Quality-Diversity (QD) have independent switches. At generation zero,
QD can allocate one validated candidate from each peer's own DIG pool.
Disabling that QD path leaves DIG's quality-first local selection intact.
Later QD uses existing PI/Chair proposals and does not call DIG or create DIG
artifacts. [Quality-Diversity Allocation](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator) owns both
allocation paths and their failure behavior.
Gems runs after measured evidence reaches a generation boundary. DIG may read
Gems as lineage/duplicate context, but cannot create or promote a Gem. New tasks
normally start with continuous evolution and enable periodic reset only after
an operator decision or diagnosis.
---
# Quality-Diversity Allocation
Quality-Diversity (QD) seeks a varied set of strong solutions rather than one
winner. The term follows Pugh, Soros, and Stanley's
[Quality Diversity: A New Frontier for Evolutionary Computation](https://doi.org/10.3389/frobt.2016.00040).
Praxist applies that principle to candidate-plan allocation, not evolutionary
genotype search.
QD is generation-aware and independently configurable. In absolute generation
0 it extends the Deep Innovation Gate (DIG) read-only planning phase. In later
generations, where DIG is off by default, it guides the existing Principal
Investigator (PI) synthesis path instead of creating another planner or
allocator artifact. Multi-PI planning adds a Chair that consolidates proposals.
The goal is to keep DIG's rigor while restoring the exploration breadth that a
direct no-DIG run can have.
## Problem
DIG improves individual peer plans:
- each peer maps the baseline before editing code;
- each peer generates several mechanism-level candidates;
- each candidate receives a structured critique;
- each implementation is locked by `selected_contract.yaml`.
The weak point is that peer-local selection can converge. If every peer sees the
same frontier, Gems, and research agenda, then many peers may select the same
obvious family: reward shaping, calibration, risk repair, or the latest Gem
lineage. This produces careful contracts, but a narrower generation.
The desired behavior is:
```text
deep individual reasoning
+ cohort-level diversity control
+ no fixed task-specific algorithm quota
+ no extra experiments during DIG
```
## Non-Goals
Initial-generation QD does not:
- change generation semantics;
- add an experiment loop;
- run training or any task-defined preliminary, aligned, or complete evaluation
during DIG;
- create a new agent runtime;
- encode task-specific metrics in Praxist core;
- force every peer into a hand-written algorithm family;
- replace Chair or PI judgment.
## Architecture
Initial-generation flow:
```text
peer context
-> DIG candidate generation
-> DIG critique
-> peer-local QD selection
-> selected_contract.yaml
-> implementation peer
```
Initial-generation QD flow:
```text
all peer contexts
-> run each peer's DIG candidate generation and critique concurrently
-> collect candidate pools and reviews
-> cohort-level QD allocator chooses one candidate per peer
-> selected_contract.yaml is updated per peer
-> implementation peers launch
```
Each peer still owns its own candidate pool. The allocator never assigns peer A's
candidate to peer B. It only decides which candidate from each peer's own DIG
pool should become that peer's locked contract.
Later-generation Multi-PI flow:
```text
completed-generation evidence
-> independent PI memos propose experiments and peer contracts
-> union of PI proposals is the candidate pool
-> Chair applies prompt-guided, soft quality-diversity allocation
-> normal research_agenda peer_contracts
-> direct implementation peers (no DIG call)
```
Later-generation single-PI flow uses that PI's normal synthesis over findings,
frontier, prior agendas, validation signals, and Gems. The PI forms proposals
and chooses the final `peer_contracts` in one existing synthesis call under the
same soft QD policy. There is no PI-memo union or Chair in this topology.
Both paths are the established non-DIG planning path with a compact policy in
prompt context. Later-generation QD is intentionally prompt-guided rather than
a second deterministic allocator: it creates no planner, contract format,
candidate file, or fact artifact. Diagnostics must therefore inspect the final
agenda and, for Multi-PI, the existing PI memos; they must not expect a
post-gen0 `dig_cohort_allocation.yaml`.
When a task declares `evaluation.diversity_dimensions`, QD-enabled PI/Chair
planning records the intended value of each applicable axis in
`peer_contracts[].planned_dimensions`. This is a plan, not experimental
evidence. The peer reports the implemented/evaluated values under the existing
finding field `design_dimensions`. Diagnostics derive the planned
Herfindahl-Hirschman Index (HHI) from the agenda and realized HHI from findings,
then report missingness and drift. Praxist does not copy plans into missing
result evidence and does not turn a missing dimension report into a hard
execution gate.
## Selection Policy
The allocator works over generic descriptors:
```text
mechanism_family
intervention_surface
intent
candidate text
risk labels
peer lane fit
local DIG selection
known frontier/Gems/sibling signatures
```
It scores candidates with:
```text
selection_score =
quality_score
+ lane_fit_bonus
+ local_selection_bonus
+ novelty_bonus
+ target_keyword_bonus
- risk_penalty
- diagnostic_penalty_when_not_in_diagnostic_slot
```
The gen0 deterministic allocator then applies cohort constraints. Later PI
synthesis receives the applicable quality, novelty, lane-fit, risk, target,
label-group, and diversity-cap controls as soft allocation guidance:
- max peers per exact diversity cell;
- max peers per mechanism family;
- max peers per intervention surface;
- max peers per intent;
- max peers sharing an intent, without assuming which task-owned intents are
diagnostic;
- optional task-defined keyword targets.
Keyword targets are generic and task-owned. Praxist core only sees named text
groups such as `architecture_or_representation` or `input_feature_use`; a task
project chooses the keywords and minimum counts.
## Target Groups
Task projects may define soft minimums under the independent policy block:
```yaml
quality_diversity:
enabled: true
initial_generation_enabled: true
later_generations_enabled: true
target_keyword_groups:
- name: architecture_or_representation
min_peers: 2
fields: [mechanism_family, intervention_surface, hypothesis, changes]
keywords: [architecture, representation, encoder, attention, model_def]
```
`initial_generation_enabled` is an independent disable switch, but gen0 QD
still needs an active gen0 DIG scope because its candidate pool comes from DIG.
Later-generation QD does not depend on DIG.
Targets are not fixed quotas. If no valid candidate matches a target, the
allocator records the miss and continues. The generation must stay live.
## Contract Construction
If the cohort allocator keeps a peer's local selected candidate, it preserves the
existing LLM-authored contract.
If the allocator chooses a different candidate from the same peer's candidate
pool, it creates a deterministic contract from the validated candidate sketch:
- `variant_name` from candidate name and peer id;
- `diversity_cell` from candidate signature;
- `mechanism_hypothesis` from candidate hypothesis;
- `files_to_modify` from candidate implementation sketch;
- `allowed_changes` from candidate changes;
- standard forbidden changes for evaluator, split, and metric calculation;
- implementation steps derived from the sketch;
- expected metric signature from the candidate diagnostic prediction;
- ablation hooks from the candidate, with a fallback hook if needed.
The normal DIG validator still checks the contract. Invalid allocations are not
allowed to reach implementation.
## Allocation Failure Behavior
QD is conservative:
- if initial-generation QD is disabled, DIG uses quality-first eligible
selection without duplicate/cell allocation;
- if later-generation QD is disabled, single-PI or PI/Chair behavior is
identical to the prior non-DIG agenda path;
- if a peer has no valid candidate alternatives, it keeps its local DIG
contract;
- if allocator validation fails for a peer, it keeps that peer's local DIG
contract and records the reason;
- if all DIG attempts fail for a peer, the existing DIG fallback-to-direct
behavior remains unchanged.
## Expected Effects
Compared with peer-local DIG, initial-generation QD should:
- preserve the stronger mechanism hypotheses and ablation discipline;
- reduce repeated same-family contracts in a generation;
- allocate at least some peers to architecture, representation, input-feature,
off-mainline, or independent exploration when candidate pools support it;
- keep diagnostic/control work bounded through the independent gen0 DIG
innovation-slot policy;
- keep Gems useful without letting every peer inherit the same Gem lineage.
Compared with no DIG, the initial-generation QD stage should:
- produce more explicit implementation contracts;
- reduce first-intuition coding;
- make failures easier to interpret;
- avoid changing task metrics, evaluator, data split, or baseline contract by
accident.
## Test Requirements
Tests should cover:
- independent DIG scope and initial/later QD config parsing;
- default DIG execution only at absolute gen0, including across Gems resets;
- cohort allocation preserves one selected candidate per peer;
- max same mechanism family is enforced when alternatives exist;
- max same intervention surface is enforced when alternatives exist;
- target keyword groups are filled when candidates exist;
- local DIG contract is preserved when allocation is disabled;
- quality-first DIG selection when initial QD is disabled;
- later single-PI synthesis and Multi-PI/Chair prompts receive QD policy only
when enabled;
- later QD does not call DIG or create a separate candidate artifact;
- deterministic override contracts pass the existing DIG validator;
- generation prompt injection uses the final cohort-selected contract;
- fallback-to-direct behavior still works after repeated DIG failure.
---
# Peer-Local Structured Memory For Long-Context Continuity
This document describes the Praxist mechanism that improves peer continuity
across multiple autonomous sessions without carrying raw transcripts forward.
## Design Goal
Each peer may span multiple runtime sessions inside one generation. A later
session should understand what the same peer already tried, what evidence it
created, what sibling peers shared, and where the current hypothesis stands. It
should not need a single unbounded chat transcript to do that.
The mechanism preserves the useful parts of a long continuous context:
- current hypothesis and open questions;
- experiment ledger and abandoned branches;
- prior session handoff;
- relevant new shared findings;
- Deep Innovation Gate (DIG) selected-contract state when present;
- anti-anchoring prompts that force reconsideration before repeating work.
It deliberately does not preserve raw message transcripts in the prompt.
## Session Boundary
The research-loop backend updates memory around each peer session:
```text
AutonomousAgentLoop._run_session()
-> build session_id
-> compose base task prompt + peer-local memory block
-> execute runtime session
-> record structured session result
```
Task-local content remains unchanged; memory supplies execution continuity and
audit artifacts only.
## Artifact Layout
For each peer:
```text
runs//gen_/peers//memory/
peer_state.yaml
experiment_ledger.jsonl
session_handoff.md
seen_shared_findings.json
memory_prompt.md
```
`peer_state.yaml` is the compact state card:
- current peer identity;
- current hypothesis;
- open questions;
- known dead ends;
- active variant;
- last session status;
- recent result artifacts discovered for the peer.
`experiment_ledger.jsonl` is the append-only local ledger:
- session id;
- concise summary;
- success/failure;
- duration and tool count where available;
- link to the session log;
- compact metrics if result artifacts are found.
`session_handoff.md` is the human-readable session boundary summary.
`seen_shared_findings.json` tracks which shared findings were already surfaced
to this peer's session prompt.
`memory_prompt.md` records the exact bounded memory block injected into the most
recent runtime prompt for auditability. Per-session prompt and manifest
snapshots are retained only up to a bounded count; the stable
`memory_prompt.md` and `session_prompt_manifest.json` files remain the latest
audit pointers.
## Prompt Injection
The runtime appends a bounded section titled:
```text
Praxist Peer-Local Structured Memory
```
The block contains:
- memory discipline requirements;
- current peer state;
- selected DIG contract snapshot if one exists;
- recent peer-local experiment ledger entries;
- new shared findings since the last session;
- previous handoff summary;
- anti-anchoring check.
The section is bounded by a character budget. If it grows too large, it is
truncated explicitly. This keeps later sessions grounded without allowing memory
to become an uncontrolled raw transcript replay.
This prompt block is peer-local and session-local. It does not broadcast full
sibling peer contracts, does not replay raw transcripts, and does not grow
linearly across generations.
## DIG Compatibility
When DIG runs, memory includes a bounded view of:
```text
runs//gen_/peers//dig/selected_contract.yaml
```
The selected-contract schema and generation scope are defined in
[Deep Innovation Gate](https://praxist.sapient.inc/en/docs/guides/deep-innovation-gate). Memory neither changes that
contract nor turns it into empirical evidence.
## Shared Findings Refresh
The memory layer reads the generation's shared-findings directory and surfaces
new JSON findings that the peer has not yet seen. This lets later sessions
benefit from sibling peers without depending on full shared transcript replay.
The prompt shows only compact metadata:
- finding id;
- finding type;
- producer;
- title or summary.
After a session ends, surfaced findings are marked as seen.
## Anti-Anchoring Behavior
Every injected memory block asks the peer to answer three questions before
continuing the same direction:
```text
1. What evidence supports continuing the current mechanism?
2. What evidence suggests pivoting, ablating, or simplifying?
3. What is the cheapest evidence that could falsify continuation?
```
This is intended to preserve the continuity benefits of long context while
reducing over-commitment to a stale local idea.
## Task Boundary
The mechanism is task-agnostic:
- it does not mention domain-specific metrics;
- it does not modify any task project;
- it does not change benchmark, evaluator, or data semantics;
- it does not alter peer count, model routing, Gems, or Frontier ranking.
Task-local prompts and role skills remain the right place for domain-specific
research instructions.
## Memory Failure Behavior
If memory files are missing, malformed, or absent, the runtime initializes a
minimal state card and proceeds. A broken memory artifact should not block a peer
from executing its assigned research task.
Session completion always attempts to write:
- a ledger row;
- an updated state card;
- a handoff note.
If the runtime session fails, the handoff and ledger still capture the failure
reason where available.
## Expected Effect
Compared with raw multi-session replay, this mechanism should:
- improve continuity across peer sessions;
- reduce repeated dead-end work;
- make session boundaries auditable;
- improve use of sibling findings;
- keep prompt growth bounded;
- preserve cross-peer diversity by forcing local anti-anchoring checks.
It is a continuity layer, not a new research selector. The generation-level
selection logic still belongs to DIG, Frontier, Gems, the Principal
Investigator (PI) panel, and its Chair.
---
# Panel Topology Prompts
A Principal Investigator (PI) agent independently proposes next-generation
work. In multi-PI topologies, a Chair compares those proposals and commits one
agenda.
> **Principle.** A `panel_topology` plugin can ship its own Jinja prompt
> templates next to its manifest. The bundled prompts under
> `praxist/plugins/workflow_stages/research_loop/backend/multi_pi/prompts/`
> are the fallback, not the only choice.
This page documents the `topology.prompts_dir` contract for panel topology
plugins. It is the override point referenced from
`PanelTopologySpec.prompts_dir` in `praxist/core/panel_topology.py`.
## Why
The multi-PI backend renders three Jinja templates:
| Template | Renderer | Purpose |
|---|---|---|
| `base.jinja2` | `BasePI.render_prompt` | Round-1 independent PI memo |
| `round2_cross_review.jinja2` | `BasePI.run_cross_review` | Round-2 anonymized cross-review |
| `chair.jinja2` | `ChairArbiter.render_prompt` | Chair synthesis prompt |
The bundled directory is the default loader source. A panel-topology plugin
that needs a different collaboration vocabulary can override one or more
templates without changing bundled files.
`topology.prompts_dir` lets a plugin ship its own prompts directory alongside
its topology contract, while keeping the bundled templates as a fallback for
everything the plugin does not override.
## Manifest field
A panel topology plugin opts in by declaring `topology.prompts_dir` in
its `plugin.yaml`:
```yaml
schema_version: 1
name: my_panel
kind: panel_topology
topology:
topology_ref: panel_topology:my_panel
prompts_dir: prompts/ # relative to this plugin directory
modes: { ... }
roles: [ ... ]
rounds: [ ... ]
```
Layout on disk:
```text
my_panel_topology/
├── plugin.yaml
└── prompts/
└── base.jinja2 # overrides bundled; chair.jinja2 etc. fall through
```
### Path resolution
`panel_topology_from_manifest(...)` (in `praxist/core/panel_topology.py`)
resolves the manifest value through `_resolve_prompts_dir`:
| `prompts_dir` value | Behavior |
|---|---|
| absent / `null` / `""` | `PanelTopologySpec.prompts_dir = None`; use bundled prompts only. |
| relative string (e.g. `prompts/`) | Resolved against the plugin directory. Must point at an existing directory. |
| absolute string | Accepted as-is. Must point at an existing directory. |
| non-string | `ValueError` at manifest time. |
| relative string without a known plugin path | `ValueError` (cannot resolve safely). |
| any value pointing at a missing directory | `ValueError` at manifest time. |
All `ValueError`s are raised at topology resolution, not at first
template lookup. Manifest-authoring bugs surface before any PI starts
rendering.
## Loader chain
When `PanelTopologySpec.prompts_dir` is supplied, both `BasePI` and
`ChairArbiter` build their Jinja `FileSystemLoader` from the search list:
```text
[plugin prompts_dir, bundled prompts dir]
```
Jinja walks the list in order, so a plugin can override one template
(say, `base.jinja2`) and let the rest fall back to the bundled version.
When `prompts_dir` is `None`, the search list contains only the bundled
directory.
This applies to all three render sites: `BasePI.render_prompt`,
`BasePI.run_cross_review`, and `ChairArbiter.render_prompt`.
## What gets threaded where
| Layer | What it does |
|---|---|
| `PanelTopologySpec.prompts_dir` | Frozen `Path \| None` on the resolved topology. |
| `panel_topology_from_manifest(..., plugin_path=...)` | Resolves the manifest value relative to the plugin directory; fails fast on missing dirs. |
| `legacy_two_round_executor.run_panel` | Resolves the topology once, pulls `topology.prompts_dir`, and threads it to both `instantiate_pi_roles(...)` and `ChairArbiter(...)`. |
| `role_bindings.instantiate_pi_roles` | Forwards `prompts_dir` to each PI constructor. |
| `BasePI.__init__` / `ChairArbiter.__init__` | Store `self.prompts_dir` and use it when building the loader. |
## Backward compatibility
A panel topology that does not declare `prompts_dir` produces
`PanelTopologySpec(prompts_dir=None)` and uses the bundled prompts. The bundled
`legacy_multi_pi_two_round/plugin.yaml` follows this path.
## Authoring checklist
When adding a panel topology plugin that ships its own prompts:
1. Place templates under `/prompts/` (or another directory
referenced by `topology.prompts_dir`).
2. Only override the templates you actually need to change. Templates
you omit fall back to the bundled versions automatically.
3. Preserve the public template variables consumed by `BasePI` and
`ChairArbiter` — overriding the layout is supported, dropping
variables silently is not.
4. Add a plugin-local unit test that asserts the override is wired
(see `tests/unit/test_panel_topology_prompts_override.py` for a
canonical pattern: render with and without the override, check for a
plugin-specific token in the output).
## See also
- `praxist/core/panel_topology.py` — `PanelTopologySpec`,
`panel_topology_from_manifest`, `_resolve_prompts_dir`.
- `praxist/plugins/workflow_stages/research_loop/backend/multi_pi/`
— `BasePI`, `ChairArbiter`, bundled prompt templates.
---
# Central Experiment Scheduler
Praxist can make one run-level scheduler the supported launch path for task
experiments. Peers submit task-defined work; they do not choose devices or create
scheduled task processes themselves.
## Why It Exists
Resource arithmetic is simple only when one component owns four facts:
1. which scientific experiment was submitted;
2. which attempt may start;
3. which process group and physical accelerator were assigned;
4. when that allocation is released.
The scheduler owns queueing, final `Popen`, process environment, infrastructure
retry, and release. Frontier promotion, evidence maturity,
[Deep Innovation Gate (DIG)](https://praxist.sapient.inc/en/docs/guides/deep-innovation-gate),
[Quality-Diversity (QD)](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator), Principal Investigator (PI)
and Chair planning, Gems, and generation policy remain separate.
## Task Contract
```yaml
compute_budget:
resource_scheduler:
mode: central
initial_concurrent_experiments: 2
min_concurrent_experiments: 1
max_concurrent_experiments: 8
supply_signal_enabled: true
supply_idle_samples: 3
supply_lease_seconds: 600
mature_supply_fraction: 0.25
mature_supply_redundancy: 3.0
mature_assessment_min_completion_probability: 0.25
exploration_reserve: 1
infrastructure_retries: 1
default_profile: gpu_work
profiles:
cpu_ordinary:
accelerator: cpu
pressure_domains: [cpu, memory, io]
gpu_work:
accelerator: gpu
gpu_count: 1
gpu_memory_gb: 20
gpu_utilization_pct: 45
pressure_domains: [cpu, memory, io]
```
Task initialization obtains these values by running the unchanged public
baseline and observing it externally. It must not create artificial CPU-only
and GPU-only rewrites. Observation is a timestamped process-lifetime series,
normally sampled every 100-200 ms. Short tasks are safely repeated or replaced
by a longer unchanged representative unit; a teardown `0%` sample or fewer
than ten useful samples does not establish zero GPU demand. Utilization uses a
robust upper estimate of full-lifetime means, while VRAM uses the observed peak
plus modest headroom. Ambiguous or undersampled demand remains unknown and
therefore exclusive. The default profile matches the public evaluator's
normal resource shape because runtime-assisted submission may omit an explicit
profile; ordinary analysis commands should not enter the experiment queue. CPU
profiles do not reserve cores per experiment; the operating system shares CPU
time and Praxist changes total experiment concurrency from live pressure. A profile
that declares CPU, memory, or I/O pressure is not newly admitted while that
declared domain exceeds its high-pressure threshold; gradual concurrency
adjustment remains the recovery path once pressure falls. GPU
profiles use two independent per-device hard limits: declared utilization plus
observed external load remains at or below 100%, and declared peak memory plus
observed external memory remains below 95% of physical VRAM. A settled job's
declared envelope is retained across setup, CPU, accelerator, evaluation, and
waiting phases because a later phase may return to its peak. Driver activity is
reported separately for diagnosis; it does not authorize transient
oversubscription. Feasible devices are ordered by the tighter of their compute
and memory headroom. Omit either GPU demand
field when demand is unknown; the profile then receives exclusive placement.
Omitted scheduler fields use documented defaults. In `mode: central`, explicit
unknown keys, misspelled booleans, malformed numbers, profiles, or pressure
domains fail during task resolution instead of silently changing scheduling
policy. Valid numeric values outside a bounded policy range retain the documented
normalization behavior.
CPU and GPU profiles are not interchangeable unless the task explicitly says
they are scientifically equivalent. Praxist never changes a failed GPU experiment
into CPU work implicitly. Omitting `--profile` selects `default_profile`; an
explicit unknown profile is rejected instead of silently using another device
class.
## Mature Evidence And Idle Supply Feedback
The scheduler adjusts capacity, but it cannot invent scientific work. When an
open generation has unused concurrency slots, the queue cannot fill them, and
consecutive host samples show headroom in task-declared pressure domains, it
writes short-lived directed leases under `gen_N/resource_supply/`. Completed
peer sessions register as idle and watch only their own lease file through the
existing event-driven loop.
The evidence controller uses the task's existing maturity policy and canonical
result store to maintain:
```text
Q = max(1, ceil(cohort_size * mature_supply_fraction))
M = unique mature results already published for this generation
D = max(0, Q - M)
A_target = min(cohort_size, ceil(mature_supply_redundancy * D))
```
When a task deliberately configures a larger hard mature close quorum, that
larger target replaces `Q` for first-wave and debt supply; otherwise the close
contract could never receive enough mature work. In that mode `M` is distinct
mature peers, matching the existing peer-quorum semantics; without a hard peer
quorum, `M` is unique mature results. Setting mature supply fraction
or redundancy to zero disables maturity-priority supply even when a hard close
quorum exists.
The inverse is also important: a positive mature supply fraction does not
create a hard close gate. Tasks that distinguish close-grade evidence must set
a positive `synthesis_trigger.mature_quorum_fraction`; otherwise raw
information density can normal-close the generation while maturity debt
remains.
Queued/running mature semantic experiments and outstanding mature-priority
leases count toward `A_target`; retries retain one semantic identity. The
default `0.25` and `3.0` values are calibrated general defaults rather than
domain truth. Setting either value to zero disables mature-priority supply
without disabling ordinary idle backfill.
During assessment, ordinary admission stops while mature top-ups remain
eligible when their calibrated probability of finishing before the generation
deadline is at least `mature_assessment_min_completion_probability`. The
compact log-normal calibration uses successful wall time divided by declared
ETA, with a neutral prior that avoids overconfidence from the first few jobs.
Before assessment, existing deadline admission remains unchanged. An unknown
ETA remains unknown rather than being converted into false precision.
Each lease names currently admissible profiles, carries `mature` or
`frontier_followup` priority, and expires if unused. `supply_lease_seconds`
sets the bounded response window (default 600 seconds, normalized to
180-3600). Expiry limits when the peer may submit an existing plan; it does not
limit the runtime of an experiment admitted before expiry. A peer may
respond with at most one already justified experiment selected by the current
research plan, evidence priorities, and exploration commitments. The signal
selects an evidence class but does not choose a hypothesis, create variants, or weaken
evaluation standards, expose a device assignment, or bypass the central queue.
The scheduler records a short-lived host-wide capacity claim so concurrent runs
cannot promise the same slot; the final launch still performs normal admission
and assigns the actual device UUID. A real queued job atomically preempts its
own or physically conflicting speculative claims, so idle feedback cannot
delay submitted research while unrelated CPU-only runs remain independent.
Once published, the claim remains stable across later pressure samples until it
is consumed, expires, is atomically preempted by real work, the generation
closes, or the run stops. Final launch always rechecks live pressure, so this
bounded response stability does not bypass admission. `supply_signal_enabled`
disables this feedback, while
`supply_idle_samples` controls its consecutive-sample requirement. Outstanding
leases count against supply capacity, so N free slots wake at most N peers. A
peer that declines an unused lease enters a bounded same-priority exponential
cooldown rather than losing eligibility for the rest of the generation. A new
experiment submission, a changed priority, or a new generation resets that
backoff; retained idle registrations allow later maturity debt to wake the peer
again without a long polling delay.
`resource_supply.stats` reports `conversion_rate=consumed/granted` and attributes
known-priority counts under `by_priority.mature` and
`by_priority.frontier_followup`. Unused
offers terminate as `declined`, `expired`, or `revoked`; a submission carrying
an already expired locator is `stale_submission`, while `reuse_ignored` is
reserved for a lease that was genuinely consumed once. These are operational
facts only and do not replace mature-result quality or quantity. Grant and
terminal transitions are durable before they affect the live lease; restart
replays the same event ledger and revokes any grant interrupted before a
response, so run-wide conversion does not reset with the scheduler process.
If host claim release is temporarily unavailable, the lease remains visible with
`release_pending: true`, is excluded from actionable maturity commitments, and is
retried by reconciliation instead of becoming hidden capacity.
At the first wave, up to `Q` peers receive direct-mature advice while at least
one peer retains exploration when the cohort has multiple peers. Once mature
commitments satisfy `A_target`, spare leases return to Pareto-relevant
follow-ups and then already planned scouts. Scientific selection remains with
the research loop.
This mechanism is resource-type neutral. CPU, memory, and I/O pressure come
from live host observations; accelerator placement remains governed by
measured memory/utilization profiles. Simulator instances, licenses, remote
services, and other bounded resources remain task-owned limits expressed in
the evaluator or through a conservative global experiment cap.
The goal is to keep the measured bottleneck supplied without forcing every
resource to a fixed utilization percentage.
## Natural-Unit Parallelism
Task harnesses should identify independent complete-evaluation units such as seeds,
folds, scenarios, simulator instances, datasets, benchmark cases, or
restart trials. A multi-accelerator profile is valid only when the evaluator
actually distributes those units across every assigned physical GPU UUID,
preserves binding through all descendants, aggregates independently of
completion order, and drains the complete process group. Declaring
`gpu_count > 1` without that implementation does not create parallelism.
Do not count the same work both as multiple top-level scheduler experiments and
as internal child units of one experiment.
Prefer the smallest unit that is independently valid, retryable, and
aggregatable without changing the scientific protocol. A long wrapper that
serially mixes setup, accelerator work, CPU evaluation, and waiting remains one
lifecycle job. Where it cannot be split safely, expose monotonic task-owned
progress. A task may fail fast after repeated identical infrastructure or
implementation errors prove the remaining units non-runnable, but must retain a
structured failure summary. Low scores, negative findings, and heterogeneous
scientific failures are not fail-fast conditions.
## Submission
```bash
PYTHONPATH="$PRAXIST_WORKSPACE_ROOT${PYTHONPATH:+:$PYTHONPATH}" \
"$PRAXIST_RUNNER_PYTHON" -m praxist.plugins.workflow_stages.research_loop.backend.protected_pids launch \
--run-dir="$PRAXIST_RUN_DIR" \
--peer="$PRAXIST_PEER_ID" \
--tag= \
--profile= \
--work-class= \
--eta= --
```
The tag identifies science, not execution syntax. Variant/protocol/data
coverage/seeds/tier belong in it; timestamps, output paths, retry numbers,
logging flags, and harmless command spelling do not. Repeated submissions with
the same identity share one queued, running, or completed job. Only exit code
75 is an automatically retryable infrastructure failure. A corrected request
whose existing job is `failed` or `rejected` must use the same scientific tag
with `--retry-terminal`; this creates a new attempt under the existing semantic
identity. An identical retransmission of an already accepted retry remains
idempotent while queued, running, completed, or `drained_unknown`; a changed
request with the flag is rejected in those states. Without the flag, a terminal
duplicate is reported explicitly instead of silently rerunning. Do not append
arbitrary retry text to a scientific tag.
Pre-launch admission timeouts and transient accelerator-probe rejections are
the exception: no experiment attempt ran, so they release the reservation and
the same scientific request may be submitted normally after capacity or host
inventory recovers.
When a peer's compatibility cap is already occupied, another central submission
from that peer remains queued until the active process group drains. Other peers
may continue to launch, and the blocked job does not consume an attempt or
create a capacity-failure record.
The final child receives an immutable attempt directory plus exact accelerator
variables. Task descendants must preserve them:
- `PRAXIST_EXPERIMENT_ID`
- `PRAXIST_EXPERIMENT_ATTEMPT_ID`
- `PRAXIST_EXPERIMENT_ATTEMPT_DIR`
- `PRAXIST_RESOURCE_PROFILE`
- `PRAXIST_ASSIGNED_GPU_UUIDS`
- `CUDA_VISIBLE_DEVICES`
- `NVIDIA_VISIBLE_DEVICES`
The built-in Linux observer inventories physical NVIDIA GPU UUIDs. It does not
advertise automatic MIG-slice discovery or placement. A task wrapper may remain
compatible with an opaque externally supplied identifier, but task
initialization must not claim that standard central scheduling validated MIG
placement.
The submitted evaluator may create ordinary worker or trainer descendants;
they inherit the same process group and resource envelope. It must not submit
each internal worker as a second top-level experiment. If a compatibility
launcher is encountered inside an active attempt, Praxist keeps that child inside
the existing allocation instead of recursively queueing it. This reuse is
verified against the run-owned attempt directory, committed READY/GO handshake,
the caller's live process group, and the scheduler's current in-memory attempt
state. Mutable attempt environment variables or copied handshake files alone do
not establish ownership.
For an active central run, `/resource_scheduler/endpoint.json` is the
run-owned launch authority. Peer shell commands may not downgrade that run to
legacy launching by changing scheduler environment variables. The environment
endpoint remains the compatibility source for legacy callers and runs that do
not have run-owned endpoint metadata.
### Optional Managed NVIDIA/CUDA Descendant Binding
This subsection applies only after the task's unchanged baseline was observed
to use, and task initialization explicitly selected, the compatible
Praxist-managed NVIDIA/CUDA backend. `PRAXIST_ASSIGNED_GPU_UUIDS` is then the authoritative ordered physical GPU
assignment. A task harness must preserve that exact value in
`CUDA_VISIBLE_DEVICES` and `NVIDIA_VISIBLE_DEVICES` across evaluator, trainer,
worker, shell, and container boundaries. A framework may use local `cuda:0`
inside the mask, but a launcher must not write that local ordinal back into a
new child's visibility environment. Missing masks may be restored from the Praxist
assignment; conflicting masks must fail clearly instead of silently rebinding.
Generated tasks using this backend should carry fast UUID, multi-UUID, missing
mask, conflicting mask, standalone, and forced-CPU contract tests. On a host
with multiple usable GPUs, launch readiness also includes a bounded non-zero
UUID parent/child CUDA check and driver-observed PID-to-UUID comparison. This is
a placement-integrity check, not a CPU/accelerator benchmark or training run.
CPU-only tasks, unified-memory systems, task-managed devices, and other
accelerator backends are valid scheduler paths and do not inherit this UUID
contract merely because an accelerator is present.
Before the evaluator executes, a small local launch barrier waits until the
semantic intent, process group, resource allocation, and protected-process
record are durable. Resume rebinds that same allocation before releasing the
barrier, so a crash cannot turn one semantic experiment into duplicate work.
## Timing And Close
Complete mature evaluations should begin early, not only after assessment
reports mature debt. `work-class=mature` has queue priority, while
`exploration_reserve` prevents mature work from consuming every slot when
exploration is queued.
When a configured mature close quorum is still missing at assessment, Praxist
stops ordinary queued/new admission but keeps deadline-safe mature top-ups
eligible. Assessment is not `CLOSING_SIGNAL`. Once the quota is met, or the
generation reaches its safety bound, the normal strict close path takes over.
At generation close, Praxist freezes the scheduler queue before writing
`CLOSING_SIGNAL`. Queued/new work is rejected; already-running process groups
continue to drain and publish evidence. Once protected work is gone, the
existing adaptive drain grace bounds agent-only cleanup and passive tool waits;
it does not kill evaluator processes. Runtime-private background-task output
files are never used as lifecycle facts because empty stdout/stderr is a valid
successful result. `praxist stop` similarly freezes all new admission before
discovering and terminating scheduler-owned process groups. The existing run
shutdown sentinel is the primary fence; a confirmed central-scheduler freeze is
an equivalent fallback. If neither succeeds, that run is left untouched and
reported in `failed_run_ids` with a nonzero CLI result. Fenced runs receive a
bounded rescan until scheduler-owned and exact run-environment descendants are
stably absent. A union bulk stop also skips its independent process-name scan
when any registry run could not be fenced, because portable hosts may not expose
enough evidence to distinguish that run's orphan from an unrelated controller.
Explicit process-scan-only operation remains unchanged. A framework-owned
`peer_workspaces/` cwd is a narrow fallback for
a descendant that cleared its environment; the run root by itself is not process
ownership. Process identities are revalidated before every signal, using the
portable `ps` start identity when procfs is unavailable, and unrelated operator
or monitor processes are not selected by broad command matching.
The live central scheduler supplies its complete active process-group set even
after a launcher exits. A legacy manifest with a live group but no verifiable
launcher is not guessed at or marked stopped; its run is returned in
`failed_run_ids` for explicit operator follow-up.
## State And Compatibility
Current state is a compact derived view at:
```text
/resource_scheduler/status.json
```
`running` is a lifecycle count: the wrapper or a descendant process group is
still alive. It is not a GPU-activity count. `running_activity` summarizes the
separate observation, and active jobs may include `resource_activity` with
`gpu_compute_active`, `gpu_context_idle`, `gpu_context_present`,
`no_gpu_process_observed`, `gpu_process_attribution_unavailable`,
`non_gpu_allocation`, or `unknown`. These fields are derived telemetry only and
never alter maturity, ranking, retries, or result validity. A single
`no_gpu_process_observed` sample can be setup, CPU work, evaluation, or a phase
transition. `gpu_process_attribution_unavailable` means the accelerator
reported a process that could not be mapped into the scheduler's PID namespace;
the accompanying `attribution` value distinguishes `complete`, `partial`, and
`unavailable` ownership mapping. `unknown` means the underlying observation was
unavailable. Admission remains conservative when ownership cannot be mapped so
shared-device external load is not mistaken for Praxist work. Use progress, logs,
process trees, and result mtimes before diagnosing a stall.
On Linux, a process group containing only zombie or exited members is terminal
work, even though `killpg(..., 0)` can still report that group as present.
Praxist reconciles that kernel state before retaining a running slot, then reaps
the wrapper and releases the allocation. A sleeping or otherwise live process
is not treated as a zombie, and platforms without procfs retain the portable
process-group check.
Attempt logs live under `/logs/experiments/`; immutable attempt metadata
lives under `/resource_scheduler/attempts/`. Existing protected-PID
manifests remain the process-lifecycle compatibility surface used by close,
stop, diagnostics, and late/quarantined result handling.
The small launch-barrier interpreter runs without Python `site` initialization
so a peer's generated `sitecustomize` cannot mistake the trusted READY handshake
for a task write. The barrier preserves the submitted environment unchanged
when it executes the real task command, so Python task descendants still load
the normal runtime guard. Peers never receive write access to scheduler state.
Tasks without `mode: central` retain the legacy launch behavior. Central mode
does not silently fall back to peer-local launching if its service cannot start.
One run-local owner lock prevents two scheduler services from controlling the
same queue. Acknowledged submissions and terminal queue rejections are fsynced
before their in-memory transitions, so restart replay neither loses accepted
work nor revives work rejected by close, freeze, or stop. Runtime environment
values that look credential-bearing by name or value shape are stored only as
hashes.
---
# Scientific Literature and Database Lookup
Praxist provides optional public literature, scientific-database, open-access,
and provenance lookup through `tool_server:literature_lookup`. It requires no
additional API key and does not change the automated experiment loop.
## Capabilities
The tool server exposes:
- `literature_search(query, sources, max_results)` for normalized cross-source
search;
- `literature_resolve(identifier)` for DOI, PMID, arXiv, or OpenAlex work
identifiers;
- `literature_open_access_text(identifier_or_url, max_chars)` for lawful open
HTML/XML text or metadata, hash, and provenance for an open PDF;
- `scientific_database_search(query, sources, max_results)` for public
scientific databases;
- `literature_source_guide(domain, objective)` for source-selection and
verification guidance.
Public sources include arXiv, OpenAlex, PubMed metadata, Crossref, Semantic
Scholar metadata, Europe PMC, UniProt, and ClinicalTrials.gov. Individual
services may rate-limit or fail. Praxist reports per-source warnings instead of
failing an unrelated research run.
## Enable in a Task
The standard tool set includes the passive lookup server. A task descriptor
should list its complete active tool set so resolve and runtime selection agree:
```yaml
praxist_plugins:
tools:
- tool_server:evaluation_tools
- tool_server:frontier_tools
- tool_server:finding_graph_query
- tool_server:memory_tools
- tool_server:prior_work_tools
- tool_server:run_report
- tool_server:literature_lookup
```
The tool is passive: network access occurs only when an agent explicitly calls
it. A task may define a task-local literature role, but Praxist core must not
contain domain-specific search strategy. Principal Investigator (PI) memo
agents and peers with the tool perform lookup; the Chair planning agent
synthesizes the evidence provided to it.
## Current-Environment-Only Rule
Search results may mention datasets, checkpoints, simulators, dependencies,
licenses, APIs, or environments that are not available locally. During a run,
agents must not acquire or install those missing resources. They should extract
the useful scientific idea and adapt it to the task's existing data, evaluator,
dependencies, hardware, and runtime.
Missing resources may be recorded as task-local limitations or future
requirements. They are not permission to mutate the host.
## Evidence and Provenance
Literature is context, not measured task performance. Records should preserve:
- title, authors, year, venue, and stable identifier;
- source and retrieval time;
- open-access status;
- exact claim supported by the source;
- uncertainty, contradiction, or negative evidence;
- a pointer that lets a later agent retrieve the original record.
A negative lookup result is also evidence. Report the query, sources attempted,
warnings, and coverage limits. Do not rewrite "not found in searched sources"
as "does not exist."
## Runtime Behavior
The lookup server:
- uses bounded timeouts and result limits;
- normalizes records without claiming cross-source identity when uncertain;
- degrades one failing source without disabling other tools;
- does not bypass paywalls or authentication controls;
- does not download task data or install dependencies;
- keeps retrieved material separate from evaluator evidence.
Use `praxist-scientific-research` when an operator agent should gather task
context before a run. Use the tool server when a running peer or PI needs a
focused source check.
---
# Troubleshooting
Use the smallest command that can identify the failing boundary. Do not edit run
artifacts to make a failed check appear healthy.
## Host or Authentication Is Not Ready
```bash
praxist doctor --json
```
For Codex-native mode:
```bash
praxist setup --profile codex-native --install-skills codex # or: claude
praxist doctor --codex-native --task-path /absolute/path/to/task --json
```
If a tested runtime package is missing or mismatched, invoke
the `praxist-runtime-install` skill instead of independently upgrading an SDK.
Codex-native authentication behavior and diagnostic-override semantics are
defined in [Credentials](https://praxist.sapient.inc/en/docs/guides/credentials#codex-native-mode-authentication).
Do not apply that repair to another setup profile.
## Package Download Certificate Failure
If installation reports `SSLCertVerificationError` or
`CERTIFICATE_VERIFY_FAILED`, repair the selected Python trust store and rerun
the pip command. On macOS, python.org distributions provide an `Install
Certificates.command` alongside the installed Python. Managed Python or
corporate environments should use their supported CA-bundle configuration.
Do not bypass TLS verification with `trusted-host`, and never place an API key
in a command argument while troubleshooting connectivity.
## Task Does Not Resolve
```bash
praxist resolve /absolute/path/to/task
```
Resolution makes no LLM calls. It reports invalid task configuration, missing
plugin descriptors, unresolved task-local references, and unsupported
agent runtime/API provider combinations before launch.
## A Started Run Disappears
`praxist start --daemonize --json` reports process creation before every
research stage necessarily initializes. Inspect:
```bash
praxist status --json
praxist --monitor --latest
```
Then read the run's `run_summary.json` and launcher log path reported by start.
A stale registry record means the registered process is no longer alive; it is
not proof that the research completed.
## Run Appears Stalled
Use:
Invoke the `praxist-diagnostic` skill in the current agent.
The diagnostic workflow separates a long active experiment from missing stage
artifacts, resource starvation, runtime/API friction, blocked generation close,
or incomplete evidence. The foreground monitor is observational and must not be
used as scientific evidence.
## Stop or Resume
Prefer the lifecycle skill:
Ask the `praxist-control` skill to stop the current run or resume the latest
run.
Interrupted generation, Principal Investigator (PI) panel, and Gems boundaries
require artifact-aware inspection before `resume`. The canonical procedure is
defined in
[Direct CLI Operations](https://praxist.sapient.inc/en/docs/guides/operators).
## Documentation Build Fails
```bash
uv sync --extra docs
uv run python scripts/build_docs_site.py
```
The build is strict. It fails on stale generated references, unowned pages,
duplicate navigation ownership, broken local links, or MkDocs warnings.
For configuration and credential details, use
[Configuration Discipline](https://praxist.sapient.inc/en/docs/concepts/config_discipline) and
[Credentials](https://praxist.sapient.inc/en/docs/guides/credentials).
---
# Platform Support
## Release Validation Matrix
| Operator host | Release status | Required verification |
|---|---|---|
| Linux on CPython 3.11 or 3.12 | Continuously tested by release CI | Run `praxist doctor` for runtime and provider readiness |
| macOS on CPython 3.11+ | Package and CLI compatibility target; not continuously tested by release CI | Run `praxist doctor` and validate all task-owned dependencies |
| Linux on other CPython 3.11+ versions | Package compatibility target; not continuously tested by release CI | Run `praxist doctor` and a task-specific smoke test |
| Windows-native | Outside the current research-runtime contract | Use a supported Linux environment instead |
The package metadata accepts CPython 3.11 or newer, but that compatibility
range must not be read as a claim that every interpreter, operating-system and
hardware combination has passed release CI. Codex or Claude Code must already
be installed and usable for skill-driven operation. Headless Linux and remote
shells are normal environments.
## Research Hardware
Praxist does not require a particular accelerator. CPU-only systems, macOS
unified memory, NVIDIA/CUDA, task-managed accelerators, and other task-owned
backends are valid when the research project itself supports them.
Task initialization observes the unchanged baseline on the current host before
declaring resource behavior. It must not infer a platform from product names,
compare artificial CPU-only and accelerator-only rewrites, or invent an
accelerator profile.
The central scheduler manages only resource classes explicitly represented by
the task harness. Its full contract is documented in
[Central Experiment Scheduler](https://praxist.sapient.inc/en/docs/guides/central-resource-scheduler).
## Out of Scope
The Praxist package and setup wizard do not provide:
- GPU drivers or CUDA;
- model-training frameworks;
- datasets or simulators;
- task-specific containers;
- cluster schedulers.
Those belong to the research project or host administrator.
`praxist resolve` can still parse task and plugin configuration without loading
POSIX locking at module import time. `praxist doctor` reports unsupported native
platforms before a research launch, and an attempted central-scheduler run fails
with a direct platform message rather than silently weakening host locking.
## Filesystem and Process Expectations
Praxist supports ordinary local paths and symlinked task or experiment storage.
The operator must have permission to create task run directories, user
configuration and registry state. The selected Python environment must be
writable by its owner.
Daemonized runs are independent of the agent conversation that launched them.
The live monitor is a separate foreground process and `Ctrl-C` exits only that
monitor.
---
# Product Usage Controls
Praxist includes optional, pseudonymized product-usage reporting. Collection is
off until the current operating-system user explicitly consents, and declining
or withdrawing does not affect installation or research. Accepting the
[Praxist User Agreement](https://praxist.sapient.inc/en/docs/legal/user-agreement) or the
[Fair Source License](https://github.com/sapientinc/praxist/blob/main/LICENSE.md)
does not enable collection.
Read the [Privacy Notice](https://praxist.sapient.inc/en/docs/legal/PRIVACY) for the complete privacy policy,
purposes, retention rules, and user rights. The
[Praxist User Data Collection Notice](https://praxist.sapient.inc/en/docs/legal/product-usage-data-notice) is
the exact versioned text presented when consent is requested. The
[Product Usage Technical Documentation](https://praxist.sapient.inc/en/docs/operations/DOCUMENTATION) describes the
implementation, endpoint, schema, storage, and failure isolation.
## Review and Choose
Review the current in-product notice and record a choice:
```bash
praxist product-usage notice
praxist product-usage consent
```
Inspect or withdraw the choice at any time:
```bash
praxist product-usage status --json
praxist product-usage withdraw
```
A plain pip installation and every non-interactive path leave consent unset.
Interactive setup presents the notice in a scrollable local terminal view.
Withdrawal stops future capture and removes unsent local events. See the
Privacy Notice for the treatment of events already delivered.
## Collector Development
Install server dependencies only on a collector development or deployment host:
```bash
uv sync --group dev --extra product-usage-server
export DATABASE_URL='postgresql+psycopg://user:password@host/database'
export COLLECTOR_INGESTION_ENABLED=true
export COLLECTOR_MAX_TABLE_BYTES=$((2 * 1024 * 1024 * 1024))
uv run alembic -c services/product_usage/alembic.ini upgrade head
uv run praxist-collector
```
Run retention in a separate process or container:
```bash
uv run praxist-retention
```
Collector ingestion is disabled until explicitly enabled. Deployment assets
and their operational controls are under `services/product_usage/`;
implementation details and audit entry points are owned by the technical
documentation.
---
# Praxist Data-Collection Module: Technical Documentation
**Last updated:** 27 August 2026
**Purpose:** As required by Section 1.4 (Data Collection) of the Fair Source License Agreement (Version 1.0), this document describes the source-code structure of the data-collection module and discloses the data-receiving endpoint URL. It is written for users, enterprise IT, and security auditors.
**Related documents:** [Privacy Notice](https://praxist.sapient.inc/en/docs/legal/PRIVACY); [Praxist User Data Collection Notice](https://praxist.sapient.inc/en/docs/legal/product-usage-data-notice) (in-product notice text, Notice version 3)
> This document describes the current implementation. Keep it and the Privacy Notice aligned with changes to the product-usage contract.
---
## 1. Module Map
| Part | Path | Notes |
| --- | --- | --- |
| Client SDK | `praxist/product_usage/` | Consent management, environment identity, event generation, local outbound queue, upload |
| Client integration | `praxist/infrastructure/product_usage.py` | Observer that projects Research Run lifecycle into telemetry events; failures are isolated from the Research Run |
| CLI entry point | `praxist/cli/product_usage.py` | The `praxist product-usage` command family (notice / consent / status / withdraw) |
| Server Collector | `praxist/product_usage/app.py`, `collector.py`, `postgres.py`, `retention.py` | HTTP ingestion, schema validation, idempotent persistence, retention deletion |
| Server deployment | `services/product_usage/` | Dockerfile, Nginx configuration, Compose, deployment scripts |
| Protocol & schema | `praxist/product_usage/protocol.py`, `schemas/v2/*.json` | Closed Schema V2 (shared by client and server) |
| Legal text | `docs/legal/product-usage-data-notice.md` | The notice text shown in-product (Notice version 3) |
## 2. Client File Map
| File | Responsibility |
| --- | --- |
| `consent.py` | Consent-state storage: a `unset` / `granted` / `denied` state machine that fails closed; atomic writes (0600 permissions); consent records are bound to the Notice version — a version mismatch counts as no consent; Agent-assisted replies recognize only `Yes` / `Agree` / `No` / `Disagree` |
| `identity.py` | Environment identity: generated at random via UUIDv4 and persisted locally in `environment.json`; never derived from any personal, device, or task information |
| `paths.py` | Fixed per-OS local file paths (Section 5); no environment-variable or project-level overrides |
| `lifecycle.py` | Generation of run-level event IDs, telemetry run IDs, event sequence numbers, and the four lifecycle events |
| `outbox.py` | Bounded local SQLite outbound queue (offline buffering, later delivery, cleared on withdrawal) |
| `batching.py` | Bounded JSON batch encoding (parsed identically on the server) |
| `transport.py` | Endpoint selection and HTTP sending (Section 3) |
| `client.py` | The `UsageSdk` facade: every collection failure is isolated and never reaches the Research Run |
| `protocol.py` | Closed Schema V2 models and boundary constants (Section 4) |
| `notice.py` | Loads the in-product notice text |
| `app.py` / `collector.py` / `postgres.py` / `retention.py` | **Server side**: HTTP entry, validation and idempotency core, PostgreSQL persistence, retention-deletion job |
## 3. Data-Receiving Endpoint URLs
| Environment | Endpoint | Notes |
| --- | --- | --- |
| **Production (release builds)** | `https://telemetry.theaiscientist.com/v1/events` | HTTPS encryption, server certificate verification, no redirect following |
| **Development (internal `.dev` builds only)** | Internal development collector (address not published) | Plain HTTP; handles only internal development/test data, never user data; not shipped with release builds |
Endpoint selection: `default_batch_sender()` in `transport.py` chooses the sender based on whether the Praxist version string contains `.dev`. The production endpoint must be a valid HTTPS URL or construction is refused outright.
Request contract:
- `POST` with `Content-Type: application/json`; request body capped at 32 KB; at most 50 events per batch;
- carries only the fixed, protocol-level User-Agent `Praxist-Product-Usage/2` — no cookies and no additional request headers;
- network timeout is 2 seconds; a success response is `202` with body `{"accepted": n, "duplicates": n}`;
- error responses: `400` (malformed request / unsupported schema version), `413` (too large), `415` (non-JSON), `503` (ingestion paused or temporarily unavailable);
- transmission failures never block or affect the Research Run; undelivered events stay in the local outbound queue and are sent automatically once connectivity returns.
## 4. Closed Schema V2 and Boundary Constants
The schema is a closed model (`extra="forbid"`): both client and server reject any out-of-schema field, and schema extensions require an explicit change to `protocol.py` plus a version bump.
| Constant | Value | Meaning |
| --- | --- | --- |
| `SCHEMA_VERSION` | 2 | Event-structure version |
| `CONSENT_NOTICE_VERSION` | 3 | Current Notice version (consent records are bound to it) |
| `MAX_BATCH_EVENTS` | 50 | Maximum events per batch |
| `MAX_REQUEST_BYTES` | 32 KB | Maximum request-body size |
| `MAX_ERROR_SUMMARIES` | 16 | Maximum error-summary groups per event |
| `MAX_ERROR_COUNT` | 65,535 | Maximum count per error group |
| `MAX_DURATION_MINUTES` | 43,200 (30 days) | Maximum recorded active-run duration |
The closed value lists for event fields and error categories are in Section 3 of the [Privacy Notice](https://praxist.sapient.inc/en/docs/legal/PRIVACY) and in `schemas/v2/usage-event.schema.json` and `schemas/v2/usage-batch.schema.json`.
## 5. Consent State and Local File Paths
All paths below are fixed, current-user-only locations with 0600/0700 permissions:
| Platform | Consent record | Outbound queue / environment identity |
| --- | --- | --- |
| Linux | `~/.config/praxist/product-usage/consent.json` | `~/.local/share/praxist/product-usage/outbox.sqlite3`, `environment.json` |
| macOS | `~/Library/Application Support/Praxist/product-usage/consent.json` | `outbox.sqlite3`, `environment.json` in the same directory |
| Windows | `%LOCALAPPDATA%\Praxist\product-usage\consent.json` | `outbox.sqlite3`, `environment.json` in the same directory |
- **Run-state files:** `runs/.json` alongside `environment.json`; stored locally only and never uploaded. **How the hash is computed** (see `run_state_path()` in `praxist/product_usage/paths.py`): the run-directory path is first normalized with `expanduser` and `resolve` (expanding the user directory and resolving it to a canonical absolute path), then UTF-8 encoded and hashed with SHA-256; the 64-character lowercase hexadecimal digest becomes the filename. SHA-256 is one-way, so the original path cannot be recovered from the filename, and since the file never leaves the machine, no path information is exposed;
- **CLI:** `praxist product-usage notice | consent | status --json | withdraw`;
- **First use:** the notice is shown and an explicit choice is awaited only in an interactive terminal while the state is `unset`; non-interactive environments remain `unset` (i.e., nothing is collected).
## 6. Offline and Failure Behavior
- **Offline or upload failure:** events are buffered in the local outbound queue and delivered when connectivity returns — **research functionality is entirely unaffected**;
- **Withdrawal (`withdraw`):** immediately stops all future capture and deletes every local unsent event;
- All client-side collection/upload exceptions are isolated by the fail-closed facade in `client.py` and never interrupt a Research Run.
## 7. Server-Side Privacy Measures
- Nginx reverse proxy: `access_log off`; strips the `X-Forwarded-For`, `X-Real-IP`, `Cookie`, and `User-Agent` headers before requests reach the application;
- Two-level rate limiting (per client and global); storage capacity ceiling `COLLECTOR_MAX_TABLE_BYTES` (default 2 GB) — new events are refused beyond it;
- Master ingestion switch `COLLECTOR_INGESTION_ENABLED`, which can pause ingestion entirely (returns 503);
- The Collector container binds to `127.0.0.1` only and reaches the managed PostgreSQL over a private endpoint; the database is never exposed to the public internet;
- `received_at` is generated by the server after validation and cannot be supplied or altered by the client;
- Retention job: runs once at service start and at least every 24 hours thereafter; events enter the deletion window on day 179 (one day of scheduling slack against the stated 180-day ceiling); if the retention job fails, it exits and the container restarts to retry;
- The server has **no** interface for deleting already-delivered events by environment identifier (a design constraint disclosed in Section 8 of the [Privacy Notice](https://praxist.sapient.inc/en/docs/legal/PRIVACY)).
## 8. How to Audit It Yourself
1. See the exact fields that would be reported: `praxist/product_usage/schemas/v2/*.json` and `protocol.py`;
2. Check current consent status: `praxist product-usage status --json`;
3. Inspect local files: the paths listed above (all readable and writable only by the current user);
4. Packet-capture verification: the endpoint, the fixed User-Agent, and the request-body content can all be verified with standard network capture tooling;
5. The complete notice text: `praxist product-usage notice`, or `docs/legal/product-usage-data-notice.md` in the repository.
## 9. Versioning and Change Management
- The schema version, Notice version, and all boundary constants are defined centrally in `praxist/product_usage/protocol.py`;
- When the notice content changes, `CONSENT_NOTICE_VERSION` increases; previously recorded consent does not carry over to the new version and is requested again at first use;
- If the implementation described in this document changes, this document, the [Privacy Notice](https://praxist.sapient.inc/en/docs/legal/PRIVACY), and the in-product notice must be updated together.
---
**Contact:** praxist@sapient.inc
---
# Agent Runtimes
Agent runtimes execute `AgentRunRequest`, emit normalized `AgentEvent` records,
and return `AgentRunResult`. The runtime controls the agent session; the API
provider (`model_provider:*`) separately describes API shape, endpoint,
model defaults, and credentials.
## Runtime Contract
Every production runtime adapter is responsible for:
- translating Praxist prompt, model, tool, cache, sandbox, and timeout intent;
- exposing selected MCP tool servers without leaking API provider response objects;
- normalizing assistant text, tool calls/results, errors, usage, and terminal
state;
- preserving cancellation and timeout status;
- redacting credentials before events enter trajectory or logs;
- supporting concurrent long-running peer sessions within its declared
capacity.
Exact capabilities vary by SDK. A runtime must report an unsupported contract
or unknown usage explicitly rather than pretending the capability exists.
## Bundled Runtimes
- `agent_runtime:claude_sdk` is the default and recommended production runtime
for new task projects, tested with `claude-agent-sdk==0.2.136`.
- `agent_runtime:codex_sdk` is an explicitly selected production runtime built
on the official `openai-codex==0.147.0` Python SDK.
- `agent_runtime:fake_runtime` is the deterministic offline runtime used by
conformance tests.
Selecting `codex_sdk` does not change the default runtime for existing tasks.
## Claude SDK Liveness
Each `agent_runtime:claude_sdk` session consumes its SDK stream on an isolated
worker loop while the research loop retains timeout and cancellation authority.
The adapter tracks complete SDK messages, partial model-stream events,
foreground tool activity, and protected background work as distinct progress
signals. Partial events are observability-only and are not copied into the
canonical agent transcript.
A liveness warning is emitted only when the observed state is `model_waiting`
and every progress source has remained silent past the warning interval. A
long-running foreground tool or active background task is reported as a
low-frequency healthy-work state instead of an SDK stream stall. Tool names may
appear in that health record, but tool inputs, commands, task IDs, and other
sensitive payloads do not. These observations do not extend or replace the
configured runtime timeout, generation deadline, stop request, or cancellation
path.
## Codex SDK Architecture
`agent_runtime:codex_sdk` uses long-lived local Codex app-server clients rather
than launching a fresh CLI command for every peer turn. An agent runtime/API
provider/credential scope may share one client while each request receives an
independent ephemeral Codex thread. The adapter consumes typed app-server
notifications and maps them to Praxist events.
Selected Praxist tool servers are attached directly as stdio MCP servers. Tool
allow/deny metadata is translated into the app-server configuration; no shell
bridge is part of the runtime contract.
The runtime also provides:
- streaming assistant, tool, reasoning, plan, file-change, usage, and terminal
notifications when the SDK emits them;
- turn interruption for Praxist stop requests and timeouts, followed by a
bounded drain;
- replacement of an unhealthy app-server client without invalidating healthy
concurrent turns prematurely;
- a private worker pool and bounded stream concurrency so large peer cohorts do
not starve unrelated Praxist async work;
- runtime-scoped state under `/runtime_state/codex_sdk/`;
- native OpenAI authentication through either `OPENAI_API_KEY` or a saved
ChatGPT login owned by the SDK-bundled Codex binary;
- a read-only account model-catalog probe so Codex-native mode launchers can
reject unsupported explicit models before starting a peer cohort;
- Praxist sandbox-intent translation for read-only, workspace-write, and full
access modes.
Codex turns keep strict caller `output_schema` values. When a known non-strict
object schema omits `additionalProperties: false`, Praxist leaves it out of the
first endpoint request instead of making a predictably rejected call. Existing
prompt parsing and task validation remain in force, and one runtime warning
records the compatibility fallback. Unrelated API provider errors and strict-schema
errors are not retried or hidden. Tasks should not shadow the bundled runtime
with model-specific schema adapters.
Usage is exact only when the app-server publishes token-usage notifications.
Otherwise the normal Praxist `usage_unknown` behavior applies. Prompt-cache
behavior remains agent runtime/API provider managed. Codex has a built-in shell surface,
so requests that require a provably shell-free runtime are rejected instead of
being represented as equivalent to Claude SDK behavior.
For Codex-native mode runs, Praxist automatically uses lossless
finding-event batching between independent threads. It keeps the complete task
contract and canonical artifacts, supplies a larger bounded set of unseen
finding references, and asks continuation sessions to reopen exact originals
only when needed. This reduces repeated thread bootstrap work without retaining
an indefinitely growing Codex conversation or relying on lossy compaction.
Stop, closing, resource-supply, timeout, and recovery semantics are unchanged.
## Capability Alignment And Differences
Both production adapters target the same Praxist request/result contract, but
they are not interchangeable implementations:
| Capability | `claude_sdk` | `codex_sdk` |
| --- | --- | --- |
| Selected Praxist MCP tools | Direct SDK MCP integration | Direct app-server stdio MCP integration |
| Streaming | Normalized from Claude SDK messages | Normalized from typed app-server notifications |
| Timeout and stop | Runtime-specific cancellation path | Turn interrupt plus bounded notification drain |
| Usage | Recorded when the SDK/API provider exposes it; otherwise unknown | Token-usage notifications when present; otherwise unknown |
| Sandbox intent | Claude SDK permission/sandbox integration | Codex read-only/workspace/full mapping; built-in shell remains part of the runtime |
| Long-run concurrency | Independent concurrent peer sessions | Independent threads over shared long-lived clients with bounded stream concurrency |
| Non-native API providers | Claude SDK/API provider compatibility path | Responses-to-Chat relay for supported Chat Completions API providers |
Do not claim complete behavioral equivalence. Tool naming, event granularity,
sandbox capabilities, cache behavior, API provider errors, and usage availability
remain SDK-specific even though Praxist normalizes their durable result shape.
## API Provider Routing
The Codex app-server consumes the Responses protocol. API provider routing is:
| API provider shape | Codex SDK path |
| --- | --- |
| OpenAI / `model_provider:openai_compatible` | Direct SDK/app-server connection |
| `model_provider:deepseek_alias` | Private run-scoped `codex-relay` to DeepSeek Chat Completions |
| `model_provider:openrouter` | Private run-scoped `codex-relay` to OpenRouter Chat Completions |
Praxist starts and stops the relay; operators must not launch a relay per peer.
The relay listens only on an ephemeral local port and receives only the selected
API provider credential. An API provider not declared compatible by the runtime plugin
must fail during resolution or startup rather than being silently rerouted.
For OpenRouter only, the relay adds a hashed run-scoped `session_id` for sticky
routing and cache locality. It does not enable response caching. DeepSeek relay
reasoning overrides are added only when the task selects an explicit policy;
`auto` preserves the existing route behavior.
Lossless context-efficiency controls are documented in
[Cost Optimization](https://praxist.sapient.inc/en/docs/guides/cost-optimization). The automatic policy applies to
Codex-native mode and OpenRouter routes; direct DeepSeek is explicitly
excluded.
## Reasoning Policy
Task projects may set one reasoning policy across API providers for every peer,
Principal Investigator (PI), Chair, and Deep Innovation Gate (DIG) planner
call:
```yaml
agent:
reasoning_effort: max # auto | off | low | high | max
```
`max` is the default for new and existing task projects that omit the field.
It requests the strongest reasoning level supported by the selected route.
`auto` is an explicit opt-in to the API provider/agent runtime's native default. `off`
explicitly disables model reasoning when the route supports that control;
`low` and `high` request the corresponding effort. The legacy
`premium_mode: true` setting remains supported as `max` when
`reasoning_effort` is `auto`; an explicit non-`auto` policy takes precedence.
For `claude_sdk` with DeepSeek, Praxist maps this policy to DeepSeek's Anthropic
compatibility fields: `thinking.type` is `enabled` or `disabled`, and enabled
requests carry the selected effort. Other Claude-compatible API providers retain
their native adaptive-thinking mapping. For `codex_sdk` with DeepSeek, the
private run-scoped relay injects the same policy into Chat Completions requests
and retains the API provider's reasoning state across tool-call subrequests.
Praxist does not summarize or reconstruct that state. Native Codex models,
including Codex-native `gpt-5.6-luna`, receive the closest supported SDK effort
(`max` maps to `xhigh`). OpenRouter's relay route uses its unified
`reasoning.effort` object.
Models whose API provider contract requires reasoning may reject `off`; Praxist
surfaces that API provider error and leaves the user's policy unchanged.
Reasoning controls change model behavior and may change latency, output-token
use, and cost. They do not change task evidence, promotion, timeout, or
generation-close contracts.
## Install And Select
Install the Codex runtime extra in the Praxist environment:
```bash
python -m pip install 'praxist[codex]'
```
The extra pins `openai-codex==0.147.0`, `claude-agent-sdk==0.2.136`, and
`codex-relay==0.5.5`, and includes the MCP dependencies used by bundled Praxist
tools. These versions are the tested runtime compatibility baseline; upgrade
them only together with Praxist runtime validation. The extra does not install a
separate task environment.
Select a runtime through a setup profile, task configuration, or explicit start
override. [Credentials](https://praxist.sapient.inc/en/docs/guides/credentials) owns API-key and saved-login setup,
authentication precedence, private Codex homes, and verification. The
[Quickstart](https://praxist.sapient.inc/en/docs/getting-started/quickstart) owns first-use profile selection;
the generated [CLI Reference](https://praxist.sapient.inc/en/docs/reference/cli) owns exact command options.
## Direct Agent Skill Use
The human-facing Codex or Claude Code CLI can invoke the bundled Praxist skills
directly. This operator host is independent of the peer runtime. A saved Codex
ChatGPT login may authenticate the official Codex SDK runtime when native
OpenAI is explicitly selected, even when Claude Code hosts the operator
workflow. Peer execution still happens through Praxist-owned runtime clients,
not by attaching to the interactive operator session.
## Adapter Checklist
When adding or changing a runtime adapter:
1. accept the runtime-neutral request/context contract;
2. translate model, prompt, MCP, sandbox, cache, timeout, and tool options;
3. emit normalized typed events and a normalized terminal result;
4. record usage when available and unknown usage otherwise;
5. redact secrets and API provider response objects;
6. preserve cancellation, timeout, and concurrent-turn isolation;
7. add offline conformance plus focused API provider/MCP integration coverage;
8. document capability differences instead of claiming cross-SDK equivalence.
---
# Open-Source Model APIs
For sustained research, Praxist generally favors APIs serving capable
open-source or open-weight models when a representative run demonstrates high
cache reuse, sufficient research quality, and stable throughput. This is a
selection policy, not a hidden runtime default: the operator still chooses the
profile during setup.
## Current Shortlist
| Priority | Option | Selection note |
|---|---|---|
| 1 | **DeepSeek V4 Pro** | Praxist provides a maintained direct API profile. Evaluate it first where the service is available and appropriate for the project. |
| 2 | **Open-source models through OpenRouter** | Use the OpenRouter profile when routing flexibility matters. Select the exact model explicitly and verify that its route reports useful cache reuse. |
| 3 | **Operator-managed open-source model endpoints** | Add or select a compatible provider plugin when deployment policy requires a private or self-hosted endpoint. Validate the plugin contract before a long run. |
The order is a practical starting point, not a claim that one model is best for
every task. Availability, pricing, model behavior, and provider-side caching can
change independently.
## Validate Before a Long Run
1. Run a short, representative workload through the intended agent runtime and
API route.
2. Confirm output quality against the task's normal evaluator rather than a
provider-specific proxy.
3. Inspect cached and uncached input usage, latency, and failures in the run
artifacts.
4. Keep the selected route only when the total cost and research quality are
acceptable together.
Praxist preserves stable prompt prefixes where the selected route supports
them, but no model name alone guarantees a high cache-hit rate. See
[Cost Optimization](https://praxist.sapient.inc/en/docs/guides/cost-optimization) for the cache contract and
[API Providers](https://praxist.sapient.inc/en/docs/guides/model-providers) for supported provider shapes.
---
# API Providers
`model_provider:*` API provider plugins describe API shape, model defaults,
credential requirements, cache capability, and route-specific compatibility.
Agent runtime plugins execute the agent loops.
## Built-In API Provider Shapes
- `model_provider:openrouter` for OpenRouter-routed model names.
- `model_provider:openai_compatible` for OpenAI-compatible endpoints.
- `model_provider:anthropic_messages` for native Anthropic Messages style.
- `model_provider:deepseek_alias` for DeepSeek-compatible aliases.
API provider names represent API format and routing. A task or operator may
override the `ModelProfile` used by a stage.
## API Provider Manifest Expectations
An API provider manifest should declare:
- supported API format;
- default model, if any;
- endpoint base, if fixed;
- required credential refs;
- cache capability;
- usage reporting capability;
- compatibility with agent runtime plugins.
## Multi-Model Runs
Research-loop agents may use different model profiles when the task contract
and selected runtime support them. Peer exploration and planning roles should
resolve providers and model names through the same configuration boundary.
Do not hard-code a task-specific model inside a generic API provider plugin.
Run-wide reasoning effort belongs to the agent runtime policy. Configure it
under `agent.reasoning_effort` as documented in
[Agent Runtimes](https://praxist.sapient.inc/en/docs/guides/agent-runtimes#reasoning-policy); adapters translate that
single policy to each API provider's supported wire contract. The default is
`max`; select `auto` explicitly to retain an API provider's native effort default.
## Provider Conformance
API provider tests should cover:
- credential resolution and redaction;
- API provider/agent runtime compatibility;
- cache capability mapping;
- missing or invalid key diagnostics;
- usage unknown behavior when the API provider does not return metering.
---
# Credentials
Credentials are resolved by Python startup code and represented by redacted
credential references. Shell wrappers do not own credential behavior.
## API Provider Credentials
Set an API key for the selected API provider before starting a run. The
maintained direct DeepSeek V4 Pro route uses `model_provider:deepseek_alias`
with `agent_runtime:claude_sdk`:
```bash
export DEEPSEEK_API_KEY=...
praxist start --model-provider model_provider:deepseek_alias \
--runtime agent_runtime:claude_sdk \
--model deepseek-v4-pro
```
See [Open-Source Model APIs](https://praxist.sapient.inc/en/docs/guides/open-source-model-apis) for route-selection
criteria. Credential handling is identical regardless of which API-backed
profile the operator selects.
When an operator explicitly selects `agent_runtime:codex_sdk`, the same
`DEEPSEEK_API_KEY` is scoped to a private run-local `codex-relay` because
DeepSeek exposes Chat Completions and the Codex app-server expects Responses.
OpenAI uses `OPENAI_API_KEY` directly without the relay. Neither path stores a
raw key in task files or human Codex CLI sessions.
### Codex-native mode authentication
Codex-native mode uses `agent_runtime:codex_sdk` with a saved ChatGPT login for
`model_provider:openai_compatible` when `OPENAI_API_KEY` is absent:
```bash
praxist setup --profile codex-native --install-skills codex
praxist start \
--codex-native \
--agent-system codex_sdk \
--model-provider model_provider:openai_compatible \
--task-path /path/to/task-project
```
The setup command verifies the Codex binary pinned inside the installed SDK,
not an unrelated executable found first on `PATH`. If that binary currently
uses API-key authentication, a local interactive terminal opens its login flow
with API provider key environment variables removed, then verifies that ChatGPT
authentication is active. A valid existing ChatGPT login is reused without a
new prompt. Noninteractive environments report the required local setup
command instead of partially configuring the profile.
This subscription check is exclusive to an explicitly selected Codex-native
profile or `--codex-native` operation. Ordinary `codex_sdk` API provider routes do
not fail readiness because of the operator's Codex login method.
`--codex-native` is authoritative after user and task configuration files are
loaded: inherited API provider/agent runtime/model defaults and API-key/custom-endpoint
variables cannot silently switch or misconfigure this run. An explicit CLI
`--model` remains authoritative. Outside this explicit mode, environment
configuration and credentials retain their normal precedence. Praxist records only a
redacted identity; for file-based login it includes a hash of the stable account
identifier, never a token. At runtime Praxist stages `auth.json`, when present,
inside a private disposable OS-temporary Codex home so the app-server can
refresh its own copy without writing the operator's `CODEX_HOME`. Keyring-backed
login uses the same private empty home and the operating-system credential
store. The private home is removed when its app-server closes and is never put
in task files, run artifacts, logs, or replay. The runtime verifies the
app-server account is actually `chatgpt` before starting a peer turn, blanks
API-key endpoint overrides, and never falls back to an API or relay when
Codex-native mode was selected.
On resume, `--codex-native` may select saved-login authentication only when the
existing run already has the canonical `codex_sdk` agent runtime and native OpenAI
API provider. Resume never rewrites a historical run's agent runtime or API provider; start a
new run to change either canonical choice.
Other API providers remain supported when the operator explicitly chooses them:
```bash
export OPENROUTER_API_KEY=...
export ANTHROPIC_API_KEY=...
```
Praxist reads `${XDG_CONFIG_HOME:-$HOME/.config}/praxist/env` by default. Use
`PRAXIST_CONFIG_FILE=/path/to/env` or command-local `--config-file
/path/to/env` for another configuration. The command-local flag takes precedence;
exported process credentials take precedence over file values.
`praxist configure-llm` manages built-in API provider configurations only. Custom
`model_provider` plugins remain supported through each plugin's documented task
or host environment contract; Praxist does not infer custom key-variable names.
Do not commit keys, paste keys into logs, or write keys into task files.
## Credential Failover Boundary
Single-key quickstart is supported. Built-in environment discovery currently
loads at most one credential per API provider. `CredentialFailoverManager` can
select a fallback when its caller supplies multiple credentials with matching
scope and API provider; an unset `target_ref` acts as a target wildcard. The
caller must also record a supported failure. Automatic runtime
failure-triggered fallback remains disabled, so this is not a user-selectable
runtime mode. Selection and failure state use redacted `CredentialRef` values.
## Tool Credentials
Tool-scoped keys are separate from API provider keys. The bundled
`tool_server:literature_lookup` is no-key-first: it uses public endpoints such
as arXiv, OpenAlex, PubMed metadata, and Crossref-style DOI metadata without
requiring task authors to configure another API provider key.
Future task-local or external plugins may add service-specific credentials for
higher rate limits or licensed sources. Those credentials are optional
enhancers, not generic Praxist requirements. Missing tool credentials must disable
or degrade only the affected lookup path and must not break tasks that do not
explicitly require that source.
## Redaction
Trajectory, logs, docs, generated sites, replay reports, and task templates must
not contain raw secrets. Tests under `tests/hardening` enforce this boundary.
---
# Cost Optimization
This guide describes low-risk token optimizations that preserve research facts
while reducing repeated context inflation.
## Goals
Praxist cost optimization follows the result-preservation principle:
- keep peer outputs and raw evidence available;
- return compact summaries by default;
- make full data available through explicit lookup;
- avoid task-specific logic in core;
- avoid asking agents to re-read large logs or broad JSON files when a small
summary is enough.
## Lossless Session Efficiency By API Provider
Praxist automatically enables lossless session efficiency for two expensive
routes:
- Codex-native mode (`agent_runtime:codex_sdk` with native OpenAI and a saved
ChatGPT login);
- `model_provider:openrouter` on any supported runtime.
Direct `model_provider:deepseek_alias` runs are always excluded. Their event
cadence, memory limits, and prompts remain unchanged, even if an operator sets
the lossless override.
The policy does not compress history or remove findings. It changes how peers
consume the same canonical artifacts:
1. stop, closing, and resource-supply events remain immediate;
2. individual shared-finding events within a short interval are collected and
followed by one continuation session;
3. the next prompt carries a larger bounded batch of unseen finding IDs plus
the existing peer state and handoff;
4. the complete task prompt remains present;
5. the continuation is told to use exact references first and reopen original
artifacts whenever details are needed or uncertain.
An already-consumed finding suppresses a duplicate wake only when both its
explicit identity and content version are unchanged. A corrected or expanded
payload with the same identity is surfaced again; missing or unparseable
identity fails open and wakes the peer.
The canonical findings, full tool outputs, session logs, and task documents
remain on disk. This is event coalescing and reference-first navigation, not
context compression.
Configuration:
```bash
# Default: auto-detect Codex-native mode and OpenRouter routes.
export PRAXIST_CONTEXT_EFFICIENCY_MODE=auto
# Change the finding-only batching interval (default 300 seconds).
export PRAXIST_CONTEXT_EFFICIENCY_MIN_SESSION_INTERVAL_SECONDS=300
# Disable finding batching for a comparison run.
export PRAXIST_CONTEXT_EFFICIENCY_MODE=off
```
`lossless` explicitly enables the policy for a non-DeepSeek route. Unknown
mode values fall back to `auto` rather than blocking a run.
When OpenRouter is selected with `agent_runtime:codex_sdk`, its existing private
relay additionally receives a non-secret, run-scoped `session_id`. This provides
sticky routing for prompt-cache locality. Praxist does not enable OpenRouter
response caching: research replies and tool calls must not be replayed verbatim.
See the
[OpenRouter prompt-caching guide](https://openrouter.ai/docs/guides/best-practices/prompt-caching)
and [response-caching distinction](https://openrouter.ai/docs/guides/features/response-caching).
Native OpenAI prompt caching remains automatic. Cache hits require exact prefix
matches, so stable instructions should precede dynamic generation/session
content. See the
[OpenAI prompt-caching guide](https://developers.openai.com/api/docs/guides/prompt-caching).
## Tool Output Limits And Full Lookup
Problem: MCP tools such as leaderboard, frontier, and finding-graph queries can
return large JSON payloads. Even when an agent only needs the top few records,
the whole response is fed back into the LLM context.
Design:
- tool handlers keep their existing business fields, such as `entries`,
`pareto_front`, `neighbor_findings`, `nodes`, and `edges`;
- handlers attach `_tool_output` metadata with:
- `schema_version`;
- `tool_name`;
- `view = summary`;
- `truncated` and `truncated_lists`;
- `full_result_ref`;
- the follow-up tool name;
- the complete JSON payload is written under the active run directory:
`tool_results/*.json`;
- agents use `mcp__evaluation-tools__read_tool_result` to read bounded chunks
of a stored result by `offset` and `max_chars`.
This gives Praxist character-window pagination without requiring a separate
cursor protocol for every tool:
```text
summary tool response -> _tool_output.full_result_ref -> read_tool_result(ref, offset, max_chars)
```
The mechanism is lossless for system-generated tool data: inline summaries may
be truncated, but the full JSON remains in the run artifacts.
Current scope:
- `evaluation-tools.get_leaderboard`;
- `frontier-tools.get_frontier`;
- `finding-graph-query.get_finding_neighbors`;
- `finding-graph-query.get_finding_subgraph`;
- `finding-graph-query.get_unlinked_recent_findings`;
- `evaluation-tools.read_tool_result`.
Out of scope:
- built-in agent CLI `Bash`, `Read`, and shell command output. Praxist cannot
hard-cap tools built into the agent runtime from a generic MCP tool server.
Prompts should still instruct agents to avoid `cat` on large JSON/log files
and to prefer compact Praxist tool responses.
## Task-Local Evaluation Toolification
Problem: expensive tasks can give peers long prompt instructions for manually
launching staged benchmarks, waiting for files, parsing benchmark JSON, applying
gates, and deciding whether to escalate. Agents often copy commands, inspect
large raw JSON files, or repeat shell checks.
Design:
- keep task-specific evaluation mechanics inside the external task project;
- provide one compact public evaluator entrypoint under the task's
`evaluations/` directory;
- keep Praxist core and generic plugins unchanged;
- have the peer implement a variant, then run one task-local command;
- preserve raw benchmark JSON and logs under the run results directory;
- print only a compact gate summary to stdout.
Example command shape:
```bash
python "$PRAXIST_TASK_PROJECT_PATH/evaluations//run.py" \
--variant-path "" \
--output-dir "/" \
--data-dir "" \
--max-stage ""
```
The task-local tool owns benchmark invocation, promotion gates, raw evidence
preservation, concise stdout summaries, and exact maturity telemetry such as
`effort_ratio` and `coverage_ratio`. For smoke runs, expose a task-owned
argument or environment variable that limits the evaluator stage without
changing Praxist core behavior.
## Measure The Effect
Savings depend on runtime caching, provider metering, evaluator output size,
event cadence, and agent behavior. Praxist does not promise a fixed token or
billing reduction.
Compare equivalent runs using canonical usage artifacts. Record input, cached
input, uncached input, output, session count, sessions per peer-generation, and
cache-hit ratio. Also verify task correctness and artifact recoverability; a
lower token count is not useful if peers repeat work or miss evidence.
Use `praxist-diagnostic` to report input, cached input, uncached input, output,
session count, sessions per peer-generation, and cache-hit ratio. A high cache
hit rate can coexist with excessive logical token use when every new thread
repeats broad bootstrap reads.
## Testing Expectations
Changes in this area should include:
- unit tests for result-ref storage, path confinement, chunk reads, and list
truncation metadata;
- tool adapter tests proving original business fields still exist;
- task-local tests for evaluator gates and reuse-existing summaries;
- offline integration tests proving startup exposes the declared tool names;
- API provider gating tests proving direct DeepSeek behavior is unchanged;
- event tests proving finding bursts coalesce while lifecycle/resource events
remain immediate;
- relay tests proving only OpenRouter receives sticky-session metadata;
- docs build after guide or docstring changes.
Do not test exact prompt prose or temporary output ordering. Test stable
contracts: bounded inline output, full-result recoverability, and task-local
gate decisions.
---
# Workflow Stages
Workflow stages are executable steps in a Praxist run.
## Research Loop
`workflow_stage:research_loop` is mandatory. It owns the peer cohort, shared
findings, frontier, finding-graph guidance, Principal Investigator (PI) and
Chair synthesis, prompt layout, generation boundaries, and run artifacts.
The stage also owns the Deep Innovation Gate (DIG), a generation-scoped
pre-code design process, and the independently configurable
Quality-Diversity (QD) allocation path. See
[Deep Innovation Gate](https://praxist.sapient.inc/en/docs/guides/deep-innovation-gate) and
[Quality-Diversity Allocation](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator).
Each generation materializes its executable topology in
`gen_/research_topology.json` before running the standard parallel peer
cohort. See [Research Topology Audit API](https://praxist.sapient.inc/en/docs/guides/research-topology-and-module-api).
## Interface Placeholders
`workflow_stage:ideation_stub` and `workflow_stage:paper_writing_stub` are
registered interface placeholders, not product modules. They remain disabled
by default and do not provide ideation or paper-writing workflows.
## Local Reviewer
`workflow_stage:reviewer_stub` provides an optional local artifact and
provenance review when explicitly run in one of these modes:
```text
local
artifact
artifacts
run_artifact
claim_check
review
```
The reviewer reads `artifact_index.jsonl`, `trajectory.jsonl`, and
`run_summary.json`, verifies artifact hashes and references, and writes
`workflow/reviewer_report.json`. It does not rerun evaluators, assess scientific
quality, or affect frontier, incubator, Gems, or leaderboard state. It refuses
to append after `run.finalized`, preserving that event as the trajectory
terminus.
## Stage Contract
An executable stage must:
- validate its input contract;
- request budget before expensive work;
- emit lifecycle events;
- write replayable artifacts;
- preserve partial outputs where safe;
- report terminal status.
Stage semantics belong in Python workflow plugins, not shell wrappers or task
harnesses.
---
# Tool Servers
Tool servers are generic plugins that expose bounded capabilities to peers and
panels. Task projects decide when a capability is scientifically relevant;
task-specific instructions do not belong in the server.
## Catalog
| Tool server | Purpose |
|---|---|
| `evaluation_tools` | Compact evaluator and leaderboard access |
| `frontier_tools` | Committed frontier views |
| `memory_tools` | Peer-memory lookup |
| `finding_graph_query` | Advisory finding-graph queries |
| `prior_work_tools` | Existing-work lookup |
| `run_report` | Human-readable derived reports |
| `literature_lookup` | Public scientific literature/database context |
| `existing_mcp_tools_shim` | Compatibility bridge for declared external tools |
The finding graph is built by a graph-maintainer plugin; its tool server is only
the query surface.
## Specialized Contracts
[Scientific Literature and Database Lookup](https://praxist.sapient.inc/en/docs/guides/scientific-literature-lookup)
owns source coverage, provenance, current-environment limits, and runtime
behavior for `literature_lookup`.
[Run Reports](https://praxist.sapient.inc/en/docs/guides/user-facing-reports-and-init) owns automatic/manual report
triggers, structure, metric interpretation, and canonical-state boundaries for
`run_report`.
## Frontier Views
`frontier_tools.get_frontier` reads committed membership from
`frontier/frontier_manifest.json`; it does not rerun promotion. The effective
task-spec maturity policy remains available for interpretation, but a live task
file cannot replace the run's frozen policy.
The latest-generation view uses current compact lanes. Historical cutoffs are
reconstructed from the canonical per-generation ledger so later capacity
eviction does not rewrite earlier membership. Compact state without immutable
history, such as unverifiable historical Gems membership, is omitted and marked
incomplete rather than guessed.
The response reports canonical and returned counts, categorized skips, policy
source, and `frontier_view_integrity_status`. If view construction hides every
entry from a non-empty canonical lane, it returns
`canonical_entries_hidden` as an integrity error. Validation candidates remain
separate and are never promoted by the reader.
## Tool Conformance
Tool-server tests cover manifest resolution, handler construction, allowed tool
names, bounded normalized output, failure normalization, redaction, and missing
optional credentials. Full payload recovery and inline output limits are
defined in [Cost Optimization](https://praxist.sapient.inc/en/docs/guides/cost-optimization#tool-output-limits-and-full-lookup).
---
# Budget Policies
Budget is dynamic in Praxist. It is not a fixed tuple copied once into a run and then
blindly enforced for every experiment.
## Budget Requests
Agents and workflow stages may request budget in the units accepted by the
current core validator:
- `tokens`;
- `wall_clock_seconds`;
- `gpu_hours`.
Requests should include scope, reason, estimated cost, expected value, and the
action that will consume the budget.
## Policy Decisions
A BudgetPolicy may:
- auto-grant low-risk requests;
- downscope a request;
- ask a Principal Investigator (PI) or Chair planning agent to review unusually
large requests;
- deny requests that would damage the run or exceed operator limits.
The default posture is result preservation. A promising experiment should be
allowed to finish when it is inside a reasonable envelope and can produce useful
artifacts, even if exact metering is imperfect.
## Usage Records
Usage records can be exact, estimated, partial, or unknown. Unknown usage must be
recorded explicitly as `usage_unknown` instead of `0`.
Late accounting failure should warn and preserve findings when possible.
## Budget Tests
Budget policy tests should cover grant, deny, downscope, review routing,
usage_unknown, replay visibility, and behavior when a peer produces results before
usage accounting completes.
---
# Agent OOBE Runbook
This runbook is the machine-facing contract for a Codex- or Claude Code-managed
Praxist installation. It keeps the agent-managed lane separate from the local
terminal wizard while reusing the same setup profiles, configuration files,
and validation.
## Boundary
Use this runbook when the operator asks Codex or Claude Code to install and
configure Praxist from PyPI, or explicitly points the agent at a source checkout
for development.
Do not use it for an ordinary runtime repair after OOBE.
An installation request does not authorize project selection or research
launch. Stop after setup and readiness; takeover is a separate action after the
operator reads [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task).
Set one conversation-local skill host, `codex` or `claude`, from the interface
running this workflow. [Agent Skills](https://praxist.sapient.inc/en/docs/user-guide/skills#install-or-refresh)
owns the installation locations and refresh commands.
The skill host is only the human interaction surface. It does not choose or
change the Praxist peer runtime selected by the setup profile.
Keep these choices in the current conversation until setup completes:
- selected Praxist executable;
- selected setup profile ID.
Do not create an OOBE state file. Legal-terms acceptance, installed configuration
(including the explicitly selected profile ID), product-usage consent,
`praxist doctor`, and task artifacts are the durable sources from which
interrupted setup is resumed.
## Interaction Contract
1. Prefer the current agent's native structured choice UI. If it is unavailable, ask the
operator to run the matching local TTY selector. Do not ask for typed
`yes`/`no` answers.
2. Use Up/Down and Enter for local choices. Treat Esc as back/cancel without
undoing a completed package installation.
3. Never ask the operator to paste an API key into chat. API providers must use
the local masked input opened by `praxist setup`; it displays one `*` per
character and keeps the key out of chat, argv, logs, and shell history.
4. Ask for a path only when bounded discovery cannot identify the intended
project. Never scan the whole home directory, a storage mount, or a dataset
tree.
5. Do not use `--force-unmanaged` automatically. If bundled skill names collide
with operator-owned paths, show the complete conflict list and offer: keep
the existing skills, back them up and replace them, or cancel. Continue only
with the operator's selected action.
6. Never accept the Praxist Fair Source License or User Agreement on the
operator's behalf or infer acceptance from an installation request. Keep
legal acceptance separate
from optional product-usage consent.
7. A usable API provider default, saved login, exported key, or successful doctor
result is not a profile choice. Only the operator may choose the profile.
## Workflow
1. **Preflight and install.** Verify Python 3.11+ and respect the operator's
active or explicitly selected Python environment. Do not replace a task
environment or modify system Python. Install the package and maintained
runtime integrations without selecting an API provider on the operator's
behalf:
```bash
python3 -m pip install --index-url https://pypi.org/simple "praxist[agents,codex]"
```
If no suitable writable environment is active, create a dedicated virtual
environment first and keep using its Python and `praxist` entrypoint for the
rest of OOBE. Do not silently install into an externally managed interpreter.
A successful pip command proves only that the distribution is installed.
Legal terms, privacy, runtime, skills, writable examples, and readiness
remain pending. Continue with the steps below; do not report
OOBE completion at this boundary. Never launch a read-only
source/package-resource example copy.
Immediately run `praxist setup --agent-managed`. This read-only command is
the machine-owned decision checkpoint. Follow its `next_required_action`
and rerun it after each completed decision.
2. **Legal terms.** Run `praxist user-agreement status --json`. If the current
legal bundle is not accepted, present three native choices with no
acceptance preselected: review the complete terms, agree and continue, or
cancel setup. For review, surface the canonical root `LICENSE.md`, the two
packaged Markdown files under `docs/legal/`, and the `license_url` and
`review_url` returned by the status command as scrollable links. Use
`praxist user-agreement review --print` only when neither interface is
available, because printing the full text into chat is the least usable
fallback. Repeat the choices after review. Explain that the Fair Source
License includes eligibility, revenue-threshold, attribution, distribution,
and use restrictions. Only after the operator explicitly agrees to the
complete bundle may the Agent run:
```bash
praxist user-agreement accept --agent-reply Agree
```
Verify that status now reports `accepted: true`. A cancellation ends OOBE
without changing runtime configuration.
3. **Privacy.** Run `praxist product-usage status --json`. Legal acceptance is
not optional telemetry consent. If collection is available and consent is
unset, offer review, share, and skip choices with neither sharing nor
skipping preselected. Review uses the packaged
`docs/legal/product-usage-data-notice.md` or its hosted page rather than
dumping the notice into chat. Record only the explicit selection with
`praxist product-usage consent --agent-reply `. If no choice is
made, leave consent unset. When `collection_available` is false, explicitly
tell the operator that this build collects no product-usage data and no
privacy authorization is required.
4. **Runtime.** Read the machine-owned options from
`praxist setup --agent-managed` (or `--list-profiles`). Present those
complete setup profiles as selectable API provider, agent runtime, concrete
model, and authentication combinations. Apply a profile only after the
operator chooses it. Before applying it, state the
profile's `authorization_detail`: Codex-native uses the saved ChatGPT/Codex
login without an API provider key or new authorization code; API-backed
setup profiles require the matching key through local masked input. Then run
`praxist setup --profile --install-skills `, where
`` is `codex` or `claude` from the current interface. Run the
command in a local interactive terminal when the profile needs either an API
key or a Codex-native ChatGPT login. The latter uses and verifies the
SDK-pinned Codex binary; no other profile may require that login. If secure
local interaction is unavailable, print the exact setup command for the
operator and wait; never route credentials through the agent conversation.
`praxist setup --interactive` handles any same-name skill conflict locally
with keep, backup-and-replace, and cancel choices. Non-interactive setup
refuses to overwrite an operator-owned path. Rerun
`praxist setup --agent-managed` and require `profile.selected: true`; a
configured API provider default without that confirmation remains incomplete.
5. **Readiness and stop.** Require `setup_decisions_complete: true`, run the
matching host diagnostics, and report the selected profile, installed skill
host, writable example locations, and any concrete blocker. The final
`next_required_action` is `run_doctor_then_finish_setup`. Do not discover or
select a research project, invoke a takeover skill, or launch a run. Link the operator to
[Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task) for the later, separate
takeover workflow.
Finish when installation, explicit setup decisions, skill registration, and
readiness checks are complete, or when one concrete blocker remains.
---
# Contributing
This repository is edited by both humans and agents. The durable rule is simple:
keep system code, generic plugins, task templates, complete examples, and real
task projects physically separate.
## Read First
- `AGENTS.md` is the machine-friendly repository contract.
- `docs/concepts/architecture.md` is the active architecture overview.
- `docs/index.md` is the source documentation entry.
## Default Change Flow
1. Identify whether the change belongs to core, a generic plugin, a task template,
a complete example, an external task project, tests, docs, or operator scripts.
2. Make the smallest change that fits the existing boundary.
3. Add or update tests at the same boundary.
4. Rebuild docs when docstrings or docs changed.
5. Record high-risk architectural implementation context in the commit or
pull-request description.
## Stable Docstrings
Public and semi-public Python APIs use Google-style docstrings. A docstring
should describe the contract that future callers can rely on, not the history of
one bug fix.
Use comments for non-obvious invariants, recovery behavior, and failure policy.
Do not add comments that only restate the next line of code.
## Default Verification
```bash
uv sync --group dev --extra docs
uv run python -m unittest discover -s tests -q
uv run python scripts/run_test_coverage.py unit --fail-under 90 --fail-under-statements 95
uv run python scripts/run_test_coverage.py integration
uv run python -m compileall -q praxist tests templates examples scripts
uv run python scripts/build_docs_site.py
git diff --check
```
Narrow changes can start with narrow tests, but the default handoff should still
include the full suite above.
The coverage command writes ignored local reports under `cover/unit/` and
`cover/integration/`. The `unit` profile is the offline non-integration test
layers and is held at 90% branch-aware total coverage plus 95% statement
coverage; the `integration` profile is `tests/integration` and remains
observational. It uses `coverage.py` from the dev dependency group and keeps the
test runner on `unittest`.
## What Requires Extra Care
Extra care is required when changing startup, plugin resolution, task path
resolution, credential selection, runtime invocation, prompt layout, event-driven
peer scheduling, finding graph guidance, budget policy, replay verification, or
run artifact schemas.
Those changes should include focused tests and, when they change architecture
contracts, a concise rationale in the pull-request description.
---
# Configuration Discipline
> Core and plugin domain code consume explicit configuration objects. Ambient
> environment variables are read at operator entry boundaries, then resolved
> once into frozen runtime configuration.
`RunConfig` centralizes run-critical settings. Narrow compatibility paths still
read ambient variables, and selected launcher or scheduler boundaries may pass
documented host values into child processes.
This contract keeps startup, replay, API provider routing, credential handling,
and agent runtime execution deterministic.
## Configuration Flow
```text
CLI entry
argparse + selected environment values
|
v
frozen RunConfig / domain contracts
|
v
core + resolved plugins
|
v
runtime-owned SDK/child-process environment
```
Configuration priority is:
```text
explicit CLI > explicit environment > override spec > task defaults
```
A value is normalized once at startup. Downstream code receives `RunConfig` or
a narrower dataclass such as `AgentRunRequest`, `ModelCallSpec`,
`CredentialRef`, or `RuntimeSandboxIntent`. It must not reconstruct API
provider, model, task, agent runtime, or credential choices from strings later in the call
graph.
## Ingress Boundaries
CLI entrypoints under `praxist/cli/` and `praxist/run.py` may read documented
Praxist configuration variables. They are responsible for:
- parsing operator intent;
- applying precedence;
- resolving task and run paths;
- selecting agent runtime/API provider/model refs;
- passing raw credentials only to credential resolution;
- producing frozen configuration for downstream consumers.
Task projects do not own research startup. The Praxist CLI resolves the task,
run directory, provider, credentials, plugins, budget, and workflow before
launch; task-specific runtime values enter through the validated task contract.
## Configuration Boundary
Code under `praxist/core/`, generic plugin domain logic, and infrastructure
services must receive configuration explicitly. Direct ambient reads are not a
transport mechanism between layers.
The following patterns are out of contract:
- reading `PRAXIST_*` values in core business logic;
- import-time constants populated from environment variables;
- mutating `os.environ` so another in-process layer can discover a value;
- inferring an API provider from model-name punctuation after startup resolved it;
- copying all host environment variables into a runtime or tool process;
- passing raw credentials through task files, prompts, logs, or trajectory.
## Runtime Egress
A runtime adapter may construct an SDK/client or child-process environment from
its explicit execution context. This is egress, not a second configuration
resolution pass.
The adapter must:
- include only the selected API provider credential;
- redact credentials from events, errors, and logs;
- pass non-secret task runtime values needed by shell or MCP children;
- avoid forwarding unrelated host secrets;
- keep runtime-private state under the selected run directory;
- preserve timeout, cancellation, and sandbox intent from the request.
For `agent_runtime:codex_sdk`, the official SDK owns the local app-server. The
runtime creates a private run-scoped relay only for supported Chat Completions
API providers and attaches selected Praxist tools directly over MCP. The
human-facing Codex CLI is a separate operator surface. When native OpenAI is
selected and no API key exists, its saved ChatGPT authentication may authorize
the SDK runtime; personal plugins, skills, hooks, MCP servers, instructions,
and runtime state do not become peer configuration.
## Credentials
Credential resolution converts a raw secret into a redacted `CredentialRef`
plus the minimum runtime material required for the selected API provider. Core
does not inspect API provider keys. An agent runtime may inject the resolved key
into an SDK or private child process when that external interface requires an
environment variable.
A runtime may also expose an optional managed-credential discovery hook. Core
accepts only a redacted `CredentialRef` whose `scope` and `provider` match and
whose `target_ref` is absent or matches the selected API provider ref.
Environment credentials win, and resolve-only startup never performs an agent
runtime authentication probe.
API provider failover, cooldown, and key selection remain Python control-plane
semantics. Shell wrappers and task projects must not duplicate them.
## Replay And Audit
Persisted startup configuration records the resolved non-secret values used by
the run. Replay and diagnostics read that frozen state rather than trying to
reconstruct the parent shell environment.
This provides:
- deterministic comparison between runs;
- one explanation for API provider/model/agent runtime selection;
- stable cost and usage attribution;
- a bounded secret-review surface;
- reproducible runtime and task-environment setup.
Unknown or unavailable values must remain explicit (`usage_unknown`, missing
optional capability, or a warning). Do not convert unknown state to a false
zero or infer it from unrelated environment variables.
## Task Experiment Configuration
Task evaluators own the scientific treatment they execute. When launch-time
arguments, environment overrides, protocol settings, or task-local config files
can change a result without changing variant code, the existing result summary
may publish a secret-free top-level `effective_config` object and an explicit
`effective_config_complete` boolean. Praxist hashes that object and propagates
only its digest, status, and the existing result-summary path; the summary
remains the single owner of the full configuration. The evaluator must publish
resolved treatment values after applying defaults and parsing, not a raw
environment snapshot. An omitted setting and an explicit setting equal to its
resolved default are the same configuration.
An evaluator claiming an exact replication should also publish the parent's
`replication_of_effective_config_sha256`. Praxist reports a match only when the
current configuration is complete and its digest equals the parent digest.
Missing provenance does not change ordinary evidence maturity, promotion,
closing, or legacy evaluator behavior; it only means the run cannot support an
exact-replication claim.
## Verification
Tests should inject configuration through constructors, CLI arguments, or a
bounded environment mapping at the entry boundary. They should also verify
that:
- serialized configuration is redacted;
- runtime child environments exclude unrelated secrets;
- API provider/model normalization happens once;
- task runtime values reach MCP/shell children through explicit context;
- replay does not depend on the ambient environment;
- exact-replication fixtures distinguish identical code under different
effective task configurations;
- new core/plugin code does not introduce undocumented environment reads.
---
# Runtime Model
This page explains the research loop at the agent-session boundary. It does not
define API provider routing, resource scheduling, evidence policy, or budgets;
those contracts have dedicated guides.
## Agent Sessions
A peer opens an agent session with a rendered prompt layout and a normalized
runtime request. The request identifies the model profile, credential reference,
tools, sandbox intent, cache policy, and timeout. The selected runtime turns SDK
events into the common Praxist event/result protocol.
After a session returns, the peer waits for a meaningful event instead of
immediately resending the same context. Examples include a revised shared
finding, lifecycle signal, resource-supply event, or heartbeat expiry. Event
cadence and continuation behavior are runtime policies; canonical findings and
task state remain on disk.
[Agent Runtimes](https://praxist.sapient.inc/en/docs/guides/agent-runtimes) owns adapter behavior and
[Cost Optimization](https://praxist.sapient.inc/en/docs/guides/cost-optimization) owns lossless finding-event
batching.
## Prompt Layout
PromptLayout V1 separates:
- **frozen blocks** shared across compatible calls;
- **semi-static blocks** changed by task or role contracts; and
- **dynamic blocks** containing generation, frontier, memory, and run state.
The rendered prompt and its layout manifest are audit snapshots. They preserve
what the agent saw for replay, but later stages regenerate views from current
canonical result, finding, frontier, Gems, memory, and boundary state.
Peer memory is a navigation index into those artifacts, not a replacement for
them. [Peer Memory](https://praxist.sapient.inc/en/docs/guides/peer-local-structured-memory-long-context)
defines its lifecycle.
## Related Contracts
- [Central Experiment Scheduler](https://praxist.sapient.inc/en/docs/guides/central-resource-scheduler):
experiment admission, process ownership, and resource release.
- [Budget Policies](https://praxist.sapient.inc/en/docs/guides/budget-policies): budget decisions and usage
records.
- [Cost Estimation](https://praxist.sapient.inc/en/docs/guides/costs): interpretation of measured usage.
- [Research Loop](https://praxist.sapient.inc/en/docs/guides/research-loop-variant-generation-flow): full
generation sequence and artifact inheritance.
---
# Generic Plugins
Generic plugins are reusable system components under `praxist/plugins/**` or another
explicit plugin root. A task project may also ship its own generic plugins under
`/.praxist/plugins/`; those are discovered as `source="task_project"` only
when `--task-path` selects the task (see
[Task Projects → Boundary Rules](https://praxist.sapient.inc/en/docs/guides/task-projects#boundary-rules)).
## Plugin Manifest
Each plugin has a `plugin.yaml` manifest describing:
- plugin kind and name;
- version and stability;
- executable entrypoint or manifest-only contract;
- dependencies and compatibility;
- declared code/assets for replay hash coverage.
The plugin loader discovers candidates, resolves dependencies, checks source
priority, and writes the selected plugin manifest into the run directory.
### Stability as an interface contract
`stability` describes the **interface contract** a plugin promises — schema,
prompt shape, role-API backward compatibility — not whether the plugin is
trustworthy to execute. Trust comes from the plugin's **source**, gated via
`TRUSTED_EXECUTION_SOURCES` (`bundled` and `task_project`).
The two paths are gated differently:
- **Bundled plugins** must declare the kind's strict expected stability
(e.g. `v1_stable` for `panel_topology`, `agent_runtime`, `workflow_stage`).
A bundled plugin affects every task project, so the contract there has to
hold.
- **Task-local plugins** under `/.praxist/plugins/` are
scope-isolated to one task project and high-churn by design; they may
declare any `stability` value (commonly `v0_experimental`) without
triggering the kind-mismatch check. Source trust is enough.
## Minimal Executable Plugin
An executable plugin usually has this shape:
```text
praxist/plugins///
plugin.yaml
adapter.py
README.md # optional, for complex plugins
```
`plugin.yaml` should declare the plugin ref, compatibility, entrypoint, and code
files that participate in source hashing. `adapter.py` should expose a small
factory or adapter object matching the kind-specific contract.
The plugin content hash covers the manifest and its declared code and assets.
Imported modules or assets omitted from the manifest are outside that plugin
content hash.
Do not name every implementation file `plugin.py` by habit. Use names that
describe the plugin's internal structure.
## Plugin-Supplied Assets
Some plugin kinds accept assets shipped alongside the manifest, declared
through dedicated manifest fields rather than ad-hoc paths.
- `panel_topology`: a plugin may declare `topology.prompts_dir` to ship
its own Jinja prompt templates for `BasePI` and `ChairArbiter`. The
bundled prompts for the multi-agent Principal Investigator (PI) panel are
used as a fallback for anything the plugin does not override. See
[Panel Topology Prompts](https://praxist.sapient.inc/en/docs/concepts/panel_topology_prompts).
When a plugin ships assets, declare the paths under `code` / `assets`
in `plugin.yaml` so they participate in replay source hashing.
## What Belongs in a Plugin
Put code in a generic plugin when it is reusable across task projects:
- runtime adapters;
- API provider adapters;
- tool servers;
- workflow stages;
- graph maintainers;
- generic budget policies.
Do not put benchmark-specific research facts or task-local role contracts into a
generic plugin. Those belong in the task project.
## Decision Test
Before adding a plugin, ask whether two unrelated task projects could use it
without copying task facts. If the answer is no, it probably belongs in a task
project.
Before adding code to core, ask whether the behavior can be selected, replaced,
or disabled through a plugin. If yes, it belongs in a plugin.
## Plugin Tests
Plugin changes should add:
- plugin-local unit tests when the plugin has real code;
- kind-specific conformance tests under `tests/conformance/`;
- workflow smoke tests when the plugin participates in startup or run execution;
- replay/hash tests when plugin code or assets affect run reproducibility.
The default test path must not require real model keys, network, GPUs, or an
external task checkout.
---
# Research Topology Audit API
Praxist records the executable research topology for each generation and
exposes a structured, read-mostly API for external modules. These surfaces are
audit and integration boundaries; they do not replace findings, frontier,
Gems, memory, or result artifacts.
## Executed Topology
The bundled executor runs one parallel cohort per generation:
```text
generation -> peer cohort -> findings -> generation boundary
-> frontier / Gems / memory / PI and Chair planning
```
Before the cohort starts, Praxist writes:
```text
gen_/research_topology.json
```
The sidecar contains a `ResearchTopologySpec` with worker nodes, edges, policy,
and metadata. It records what the executor is about to run. Generic worker
types in the schema are descriptive vocabulary; the bundled executor does not
claim to execute a worker type unless that node appears in the materialized
topology.
## Module API
`ResearchLoopModuleAPI` provides these operations:
```text
submit_recommendation
request_topology_change
list_commands
list_findings
get_frontier_summary
get_validation_signals
get_gems_summary
get_memory_summary
get_run_status
```
`get_frontier_summary` returns durable frontier evidence.
`get_validation_signals` returns compact task-defined signals for triage and
follow-up planning. Validation signals retain their actual stage and coverage;
they do not become clean parents unless the task contract explicitly grants
that authority.
## Command Queue
Recommendations and topology-change requests are appended to:
```text
external_requests/research_commands.jsonl
```
The bundled executor records these commands but does not inject them into peer
prompts or mutate a live topology. A queued command is operator intent, not an
experiment contract. PI and Chair planning continue to own executable peer
assignments.
```text
external module -> command queue -> audit and operator review
```
External modules should use this API instead of editing prompts, agendas, or
generation artifacts directly.
## Boundary Rules
- `GenerationLoop` owns lifecycle, resume, and generation boundaries.
- The topology executor owns cohort execution behind that boundary.
- Topology policy and worker adapters belong in plugins, not task-specific core
branches.
- Requests use the current `queue_for_generation_boundary` policy.
- Queued requests have no scientific or promotion authority by themselves.
---
# CLI Reference
This page is generated from the same `argparse` tree used by the
`praxist` executable. The command implementation is the sole definition;
rebuild the documentation after changing CLI arguments.
```text
usage: praxist [-h] [--version] ...
```
## Global arguments
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
| `--version` | no | show program's version number and exit |
## Commands
| Command | Purpose |
|---|---|
| [`praxist configure-llm`](#praxist-configure-llm) | Persist a built-in Praxist LLM provider profile. |
| [`praxist docs`](#praxist-docs) | Open or print the hosted Praxist documentation. |
| [`praxist doctor`](#praxist-doctor) | Check Praxist host readiness. |
| [`praxist examples`](#praxist-examples) | List or install complete writable example projects. |
| [`praxist install-skills`](#praxist-install-skills) | Install bundled Praxist skills for Codex or Claude Code. |
| [`praxist uninstall-skills`](#praxist-uninstall-skills) | Remove Praxist-managed agent skill registrations. |
| [`praxist product-usage`](#praxist-product-usage) | Review or change pseudonymous product-usage consent. |
| [`praxist monitor`](#praxist-monitor) | Watch Praxist run state in a live read-only terminal dashboard. |
| [`praxist resolve`](#praxist-resolve) | Resolve a task project's plugin manifest (no LLM calls). |
| [`praxist resume`](#praxist-resume) | Resume an interrupted Praxist run. |
| [`praxist setup`](#praxist-setup) | Configure this host for Praxist operation. |
| [`praxist start`](#praxist-start) | Launch a new Praxist research run (registry-backed). |
| [`praxist status`](#praxist-status) | List known Praxist experiment runs. |
| [`praxist stop`](#praxist-stop) | Stop a Praxist run by run_id, or stop everything with --all. |
| [`praxist takeover`](#praxist-takeover) | Open Codex or Claude Code and hand off a project to Praxist takeover. |
| [`praxist uninstall`](#praxist-uninstall) | Remove the user-level Praxist installation. |
| [`praxist user-agreement`](#praxist-user-agreement) | Review the Praxist License and User Agreement or inspect acceptance status. |
## `praxist configure-llm`
Persist a built-in Praxist LLM provider profile.
```text
usage: praxist configure-llm [-h] --provider PROVIDER [--model MODEL]
[--agent-system {claude_sdk,codex_sdk}]
[--api-key-stdin | --api-key-env API_KEY_ENV | --no-api-key | --remove-api-key]
[--config-file CONFIG_FILE] [--project-env-file PROJECT_ENV_FILE]
[--print-source-command] [--json] [--dry-run]
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
| `--provider` | yes | Built-in provider name or compatible provider plugin reference. |
| `--model` | no | Provider model name to persist. |
| `--agent-system` | no | Agent runtime selection to persist. Choices: `claude_sdk`, `codex_sdk`. |
| `--api-key-stdin` | no | Read the provider API key from stdin; a local TTY shows one * per character. |
| `--api-key-env` | no | Read the provider API key from this environment variable. |
| `--no-api-key` | no | Update non-secret provider settings without writing an API key. |
| `--remove-api-key` | no | Remove this provider's stored API key from the selected config file(s). |
| `--config-file` | no | Config file to update (default: $PRAXIST_CONFIG_FILE or the user config). |
| `--project-env-file` | no | Also write Praxist LLM config to this explicit task-local .env file. |
| `--print-source-command` | no | Print the shell command that loads the selected config file. |
| `--json` | no | Emit the result as JSON. |
| `--dry-run` | no | Validate and report changes without writing files. |
## `praxist docs`
Open or print the hosted Praxist documentation.
```text
usage: praxist docs [-h] [--no-open]
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
| `--no-open` | no | Print the documentation URL without opening a browser. |
## `praxist doctor`
Check Praxist host readiness.
```text
usage: praxist doctor [-h] [--json] [--task-path TASK_PATH] [--config-file CONFIG_FILE]
[--agent-system {claude_sdk,codex_sdk}] [--model-provider MODEL_PROVIDER]
[--model MODEL] [--codex-native] [--target {auto,codex,claude}] [--advisory]
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
| `--json` | no | Emit the readiness report as JSON. |
| `--task-path` | no | Also validate this task project and its runtime environment. |
| `--config-file` | no | Config file to inspect (default: $PRAXIST_CONFIG_FILE or the user config). |
| `--agent-system` | no | Check one research runtime (default: configured runtime or claude_sdk). Choices: `claude_sdk`, `codex_sdk`. |
| `--model-provider` | no | Check one model_provider ref using the same precedence as praxist start. |
| `--model` | no | Check this selected model (Codex-native verifies it in the account catalog). |
| `--codex-native` | no | Check codex_sdk with native OpenAI and the saved ChatGPT login. |
| `--target` | no | Check bundled skills for this agent host (default: detect managed installs). Choices: `auto`, `codex`, `claude`. Default: `auto`. |
| `--advisory` | no | Always return exit 0 while retaining readiness failures in the report. |
## `praxist examples`
List or install complete writable example projects.
```text
usage: praxist examples [-h] ...
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
## `praxist install-skills`
Install bundled Praxist skills for Codex or Claude Code.
```text
usage: praxist install-skills [-h] [--target {codex,claude}] [--target-dir TARGET_DIR]
[--mode {copy,symlink}] [--replace] [--force-unmanaged]
[--migrate-legacy-symlinks] [--dry-run] [--json]
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
| `--target` | no | Skill host. Default: codex. Choices: `codex`, `claude`. Default: `codex`. |
| `--target-dir` | no | Override the target skill directory. |
| `--mode` | no | Register skills by copying package content or linking a source checkout. Choices: `copy`, `symlink`. Default: `copy`. |
| `--replace` | no | Refresh existing Praxist-managed entries; unmanaged paths require --force-unmanaged. |
| `--force-unmanaged` | no | With --replace, back up and replace unmanaged entries whose names exactly match bundled Praxist skills. Unrelated skills are untouched. |
| `--migrate-legacy-symlinks` | no | With --replace, explicitly adopt old Praxist repo-style symlinks that predate the ownership manifest. |
| `--dry-run` | no | Report actions without changing the target directory. |
| `--json` | no | Emit the result as JSON. |
## `praxist uninstall-skills`
Remove Praxist-managed agent skill registrations.
```text
usage: praxist uninstall-skills [-h] [--target {codex,claude}] [--target-dir TARGET_DIR] [--dry-run]
[--json]
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
| `--target` | no | Skill host. Default: codex. Choices: `codex`, `claude`. Default: `codex`. |
| `--target-dir` | no | Override the target skill directory. |
| `--dry-run` | no | Report removals without changing the target directory. |
| `--json` | no | Emit the result as JSON. |
## `praxist product-usage`
Review or change pseudonymous V2 product-usage consent. Withdrawal stops future capture and deletes unsent local events; delivered events expire through scheduled retention.
```text
usage: praxist product-usage [-h] ...
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
## `praxist monitor`
Render a live read-only dashboard from praxist status, orchestrator snapshots, peer memory health, recent logs, and lightweight host load. The dashboard runs directly in the current terminal and never controls the Praxist research process.
```text
usage: praxist --monitor [-h] [--run-id RUN_ID] [--run-dir RUN_DIR] [--task-path TASK_PATH]
[--latest] [--interval INTERVAL] [--once] [--follow] [--no-clear] [--plain]
[--log-lines LOG_LINES] [--peer-limit PEER_LIMIT]
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
| `--run-id` | no | Monitor one run id. |
| `--run-dir` | no | Monitor one run dir. |
| `--task-path` | no | Prefer active rows for this task path. |
| `--latest` | no | Select the latest active run row when more than one exists. |
| `--interval` | no | Frame interval in seconds (default: 0.2 for the fullscreen TUI, 1 for plain text). |
| `--once` | no | Render one frame and exit. |
| `--follow` | no | Keep refreshing even when stdout is not an interactive terminal. |
| `--no-clear` | no | Append frames instead of clearing the terminal between refreshes. |
| `--plain` | no | Use the legacy plain-text monitor instead of the fullscreen TUI. |
| `--log-lines` | no | Recent log lines to show for the selected run (default: 18). Default: `18`. |
| `--peer-limit` | no | Maximum peer rows to render (default: 24). Default: `24`. |
## `praxist resolve`
Discover and resolve a task project's plugin manifest without making any LLM calls. Equivalent to:
python -m praxist.run run --task-path --resolve-only --local
Exits non-zero on resolution failure (manifest schema error, missing plugin, etc.). On success, emits a JSON document on stdout summarizing the resolved run identity.
```text
usage: praxist resolve [-h] [--config-file CONFIG_FILE] [--agent-system {claude_sdk,codex_sdk}]
[--workspace WORKSPACE] [--run-dir RUN_DIR] [--runtime RUNTIME]
[--codex-native] [--model-provider MODEL_PROVIDER]
[--budget-policy BUDGET_POLICY] [--credential-profile CREDENTIAL_PROFILE]
[--model MODEL] [--result-summary RESULT_SUMMARY]
[task_path]
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
| `task_path` | no | Path to the task project directory (default: current directory). Default: `.`. |
| `--config-file` | no | Config file to load (default: $PRAXIST_CONFIG_FILE or the user config). |
| `--agent-system` | no | Agent system used to resolve runtime/provider defaults. Choices: `claude_sdk`, `codex_sdk`. |
| `--workspace` | no | Workspace directory (default: current working directory). Default: ``. |
| `--run-dir` | no | Override run artifact directory. Defaults to the task project's runtime_outputs.root / experiments directory; paths inside the Praxist source checkout are rejected. Default: ``. |
| `--runtime` | no | Override agent_runtime plugin ref. Default: ``. |
| `--codex-native` | no | Resolve with codex_sdk, native OpenAI, and saved ChatGPT login while ignoring API-key/custom-endpoint configuration. |
| `--model-provider` | no | Override model_provider plugin ref. Default: ``. |
| `--budget-policy` | no | Override budget_policy plugin ref. Default: ``. |
| `--credential-profile` | no | Override credential profile name (rarely needed for resolve-only). Default: ``. |
| `--model` | no | Override agent model. Not used by resolve-only itself, but propagated for parity. Default: ``. |
| `--result-summary` | no | Validate one evaluator-produced JSON summary against the task's maturity telemetry contract before resolving. Default: ``. |
## `praxist resume`
Continue an existing Praxist run directory from its last safe completed generation boundary. The target may be a registry run_id from ``praxist status`` or a direct experiments/run_* path.
```text
usage: praxist resume [-h] [--task-path TASK_PATH] [--config-file CONFIG_FILE]
[--agent-system {claude_sdk,codex_sdk}] [--runtime RUNTIME_REF]
[--codex-native] [--model MODEL] [--model-provider MODEL_PROVIDER_REF]
[--strategy {auto,mixed,explore,exploit}] [--cohort COHORT]
[--generations GENERATIONS] [--server] [--daemonize]
[--resume-policy {completed_generation}] [--force]
[--startup-timeout STARTUP_TIMEOUT] [--json]
target
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
| `target` | yes | Registry run_id or path to an existing Praxist run directory. |
| `--task-path` | no | Override task project path when resuming from a run directory. |
| `--config-file` | no | Config file to load (default: $PRAXIST_CONFIG_FILE or the user config). |
| `--agent-system` | no | Override agent system for the resumed launch. Choices: `claude_sdk`, `codex_sdk`. |
| `--runtime` | no | Override agent_runtime plugin ref. |
| `--codex-native` | no | Resume in Codex-native saved-login mode without provider API keys. |
| `--model` | no | Override model name. |
| `--model-provider` | no | Override model_provider plugin ref. |
| `--strategy` | no | Override frontier strategy. Choices: `auto`, `mixed`, `explore`, `exploit`. |
| `--cohort` | no | Cohort size override (exported as COHORT_SIZE). |
| `--generations` | no | Maximum generations override (exported as MAX_GENERATIONS). |
| `--server` | no | Disable --local mode (server mode). |
| `--daemonize` | no | Use the same double-fork daemon launch path as praxist start. |
| `--resume-policy` | no | Resume policy forwarded to praxist.run. Choices: `completed_generation`. Default: `completed_generation`. |
| `--force` | no | Allow resume only when an old registry entry's process ownership cannot be verified. It never overrides a verified live controller. |
| `--startup-timeout` | no | Seconds to wait for resume startup artifacts. Default: `30.0`. |
| `--json` | no | Emit one JSON document on stdout instead of the operator summary. |
## `praxist setup`
Pip-first Praxist host setup. Run this after installing the package and runtime extras. It writes only Praxist user-level configuration and Praxist-managed agent skill registrations; it does not install global agent CLIs or task-specific dependencies.
```text
usage: praxist setup [-h] [--agent-system {claude_sdk,codex_sdk}] [--provider PROVIDER]
[--model MODEL] [--api-key-stdin | --api-key-env API_KEY_ENV | --no-api-key]
[--interactive]
[--profile {codex-native,deepseek-api,openrouter-api,anthropic-api}]
[--list-profiles] [--agent-managed] [--install-skills {codex,claude,none}]
[--config-file CONFIG_FILE] [--json] [--dry-run] [--skip-doctor]
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
| `--agent-system` | no | Agent runtime selection to persist. Choices: `claude_sdk`, `codex_sdk`. |
| `--provider` | no | Built-in provider name to configure. |
| `--model` | no | Provider model name to persist. |
| `--api-key-stdin` | no | Read the provider API key from stdin; a local TTY shows one * per character. |
| `--api-key-env` | no | Read the provider API key from this environment variable. |
| `--no-api-key` | no | Configure a supported no-key authentication route. |
| `--interactive` | no | Review the License and User Agreement, choose optional privacy, and select a coherent runtime profile in a local TTY wizard. |
| `--profile` | no | Apply one complete profile; a missing API key is requested in the local terminal. Choices: `codex-native`, `deepseek-api`, `openrouter-api`, `anthropic-api`. |
| `--list-profiles` | no | List supported setup profiles as JSON and exit without changes. |
| `--agent-managed`, `--codex-managed` | no | Report the read-only agent-managed first-use decision state and next required action as JSON. --codex-managed remains a compatibility alias. |
| `--install-skills` | no | Install bundled skills for an agent host (interactive default: codex). Choices: `codex`, `claude`, `none`. |
| `--config-file` | no | Config file to update (default: $PRAXIST_CONFIG_FILE or the user config). |
| `--json` | no | Emit setup and readiness results as JSON. |
| `--dry-run` | no | Validate and report changes without writing files. |
| `--skip-doctor` | no | Skip the final readiness report. |
## `praxist start`
Async launcher: starts ``python -m praxist.run run`` in a new session, redirects stdout/stderr to a run-local log file, and writes a registry entry under $PRAXIST_STATE_DIR/runs/.
Pass --task-path / --model / --model-provider to override the resolved task and runtime configuration.
```text
usage: praxist start [-h] [--task-path TASK_PATH] [--config-file CONFIG_FILE]
[--agent-system {claude_sdk,codex_sdk}] [--runtime RUNTIME_REF]
[--codex-native] [--run-dir RUN_DIR] [--resume] [--resume-from RESUME_FROM]
[--resume-policy {completed_generation}] [--model MODEL]
[--model-provider MODEL_PROVIDER_REF] [--strategy {auto,mixed,explore,exploit}]
[--cohort COHORT] [--generations GENERATIONS] [--server] [--daemonize]
[--startup-timeout STARTUP_TIMEOUT] [--json]
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
| `--task-path` | no | Task project directory (default: $TASK_PATH or the current directory). |
| `--config-file` | no | Config file to load (default: $PRAXIST_CONFIG_FILE or the user config). |
| `--agent-system` | no | Agent system the launched run will use. Default: $PRAXIST_AGENT_SYSTEM if set, else 'claude_sdk'. Recognised values: claude_sdk (default), codex_sdk. Choices: `claude_sdk`, `codex_sdk`. |
| `--runtime` | no | Explicit ``agent_runtime:*`` plugin ref. Wins over the agent-system mapping when set. |
| `--codex-native` | no | Use codex_sdk with native OpenAI and saved ChatGPT login, ignoring API-key and custom-endpoint settings from process/config/task env. |
| `--run-dir` | no | Explicit run directory (default: /experiments/run__). |
| `--resume` | no | Resume an existing run directory instead of requiring fresh artifacts. |
| `--resume-from` | no | Path to an existing run directory to resume. Equivalent to --run-dir --resume. |
| `--resume-policy` | no | Resume policy forwarded to praxist.run. Choices: `completed_generation`. Default: `completed_generation`. |
| `--model` | no | Model name forwarded to the runtime; defaults depend on provider. |
| `--model-provider` | no | Provider plugin ref (e.g. model_provider:deepseek_alias). Default cascades from agent system: claude_sdk → deepseek_alias when DEEPSEEK_API_KEY is set, then openrouter when OPENROUTER_API_KEY is set, then anthropic_messages; codex_sdk follows the same credential-aware selection and falls back to openai_compatible. |
| `--strategy` | no | Frontier strategy (auto\|mixed\|explore\|exploit). Choices: `auto`, `mixed`, `explore`, `exploit`. Default: `auto`. |
| `--cohort` | no | Cohort size override (exported as COHORT_SIZE). |
| `--generations` | no | Maximum generations override (exported as MAX_GENERATIONS). |
| `--server` | no | Disable --local mode (server mode). |
| `--daemonize` | no | Double-fork the launcher before spawning so the workload survives when the launching shell's process tree is reaped. Required for sandboxed launcher contexts (agent tool shells, CI runners, Docker ``--init``). The default ``start_new_session=True`` path is fine for a normal terminal. |
| `--startup-timeout` | no | Seconds to wait for startup artifacts before returning. A live run that exceeds the deadline remains in 'starting' state (default 30). Default: `30.0`. |
| `--json` | no | Emit one JSON document on stdout instead of the operator table. |
## `praxist status`
Merge the run registry written by ``praxist start`` with a cross-platform ``ps`` scan to list every Praxist run the operator should know about.
Rows are tagged with their source: ``registry`` (managed run, PID alive), ``ps-only`` (matching process without a registry entry — e.g. started through a direct Python invocation), or ``stale`` (registry entry whose PID is gone).
```text
usage: praxist status [-h] [--json] [--run-id RUN_ID] [--task-path TASK_PATH] [--active] [--latest]
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
| `--json` | no | Emit one JSON document on stdout instead of the plain-text table. |
| `--run-id` | no | Show only this registry run id. |
| `--task-path` | no | Show runs for this task directory. |
| `--active` | no | Show only live local runs. |
| `--latest` | no | Show only the newest matching run. |
## `praxist stop`
``praxist stop `` terminates one specific run via its registry entry. ``praxist stop --all`` terminates every Praxist-recognised process — by default the union of registry entries and ``ps``-scan matches.
Registry-backed runs close new admission before discovery. Both modes send SIGTERM, wait --grace seconds, then SIGKILL any process still alive; registry-backed runs also perform a bounded stable-empty rescan for late children.
```text
usage: praxist stop [-h] [--all] [--registry-only] [--ps-scan-only] [--grace GRACE_SECONDS] [--gc]
[--dry-run] [--json]
[run_id]
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
| `run_id` | no | Run id (filename stem of $PRAXIST_STATE_DIR/runs/.json). |
| `--all` | no | Stop every recognised Praxist run (registry + ps-scan by default). |
| `--registry-only` | no | With --all: only target registry-managed runs. |
| `--ps-scan-only` | no | With --all: only target unregistered runs found by the process scan. |
| `--grace` | no | Seconds to wait after SIGTERM before SIGKILL (default 5.0). Default: `5.0`. |
| `--gc` | no | Remove stale registry entries. A stale entry is one whose recorded PID is no longer alive, or whose live command line no longer matches the prefix recorded at ``praxist start`` time (PID recycling). No signals are sent. |
| `--dry-run` | no | Show what would be signalled without sending any signals. With --gc, list the would-be-removed entries without deleting any files. |
| `--json` | no | Emit a JSON outcome document instead of the operator summary. |
## `praxist takeover`
Open Codex or Claude Code and hand off a project to Praxist takeover.
```text
usage: praxist takeover [-h] [--task-path TASK_PATH] [--codex-native | --configured-provider]
[--operator {codex,claude}] [--yes] [--dry-run] [--json]
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
| `--task-path` | no | Research project to hand off (default: select locally or use the current directory). |
| `--codex-native` | no | Use the no-key Codex-native takeover skill. |
| `--configured-provider` | no | Use the configured-provider takeover skill. |
| `--operator` | no | Agent CLI that hosts the takeover workflow. Default: codex. Choices: `codex`, `claude`. Default: `codex`. |
| `--yes` | no | Launch without the final Enter confirmation. |
| `--dry-run` | no | Show the redacted handoff without starting the agent CLI. |
| `--json` | no | Emit the redacted handoff as JSON. |
## `praxist uninstall`
Remove Praxist-managed CLI files, runtime environment, agent skills, configuration, state, and cache. Research projects, task environments, run directories, agent CLIs, Python, and uv are never removed.
```text
usage: praxist uninstall [-h] [--venv-dir VENV_DIR] [--bin-dir BIN_DIR] [--skills-dir SKILLS_DIR]
[--keep-user-data] [--dry-run] [--json]
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
| `--venv-dir` | no | Override the Praxist-managed virtualenv path. |
| `--bin-dir` | no | Override the user bin directory containing Praxist entrypoints. |
| `--skills-dir` | no | Override skill removal with one explicit managed directory. |
| `--keep-user-data` | no | Keep Praxist configuration, registry state, product-usage state, and cache. |
| `--dry-run` | no | Validate ownership and report removals without changing files. |
| `--json` | no | Emit one machine-readable result document. |
## `praxist user-agreement`
Review the Praxist License and User Agreement or inspect acceptance status.
```text
usage: praxist user-agreement [-h] ...
```
| Argument | Required | Description |
|---|---:|---|
| `-h`, `--help` | no | show this help message and exit |
---
# Skills Reference
This catalog is generated from bundled `SKILL.md` front matter. Each skill
file is the sole definition of its activation contract and workflow.
Mechanism abbreviations use the compact definitions in the
[Glossary](https://praxist.sapient.inc/en/docs/about/glossary).
| Skill | Invoke in Codex | Purpose |
|---|---|---|
| `praxist-control` | `$praxist-control` | Start, stop, resume, inspect status, open the independent read-only foreground TUI, detect active runs, and safely repair Praxist run lifecycle boundaries from a supported agent interface. Use when the user asks to launch a Praxist task, stop or kill a run, continue/resume an interrupted run, restart from the latest safe generation, query current run progress, open or exit a live monitor, detect or list currently running Praxist tasks in the environment, inspect generation status, view incubator/frontier/leaderboard performance, check hardware load, handle interrupted PI panel or Gems reset boundaries, inspect whether a task directory is runnable, or control Praxist lifecycle commands with `praxist start`, `praxist stop`, `praxist status`, `praxist --monitor`, `praxist resume`, or `praxist resolve`.
Source: `skills/praxist-control/SKILL.md` |
| `praxist-diagnostic` | `$praxist-diagnostic` | Diagnose Praxist run health, artifact integrity, research-loop completeness, generation-scoped DIG/QD, PI/Gems/frontier/incubator consistency, peer memory freshness, diversity HHI, hardware utilization, LLM/runtime friction, task harness health, sustained low-performance causes, strongest variants/Pareto front, strong-variant lineage, and human-readable run reports. Use when the user asks an agent to investigate whether a current or historical Praxist run is healthy, why progress or performance is weak, whether artifacts or promotions are missing, whether guard or resource issues are blocking peers, to produce a detailed agent behavior analysis report, to generate an A/B/C run report, or to improve or optimize a task after diagnosis using task-directory-only parameter and prompt changes. Default diagnostics are analysis-only; explicit improvement mode may stop the selected run and edit task-level configuration/prompts, but must not modify Praxist core logic.
Source: `skills/praxist-diagnostic/SKILL.md` |
| `praxist-interactive-task-init` | `$praxist-interactive-task-init` | Build a Praxist task project through a confirmation-first interactive agent workflow. Use when a user wants Praxist task initialization with human confirmation of research goals, constraints, metrics, ranking rules, evaluation protocol, compute budget, baseline handling, or launch readiness; when the user asks for an interactive task init skill; or when the agent should propose a task harness first and ask the user to approve or revise it before writing files.
Source: `skills/praxist-interactive-task-init/SKILL.md` |
| `praxist-onboarding` | `$praxist-onboarding` | Establish detailed context for Praxist before helping a user install, configure, operate, troubleshoot, or extend the system. Use when the user is new to Praxist, has installed or is about to install the package, asks what Praxist does, asks how Praxist works, asks about `praxist` commands, task projects, runs, peers, generations, frontier/incubator/Gems, PI panels, configuration, API keys, model providers, agent runtimes, plugin architecture, run artifacts, software boundaries, or asks an agent to inspect whether the local environment is ready for Praxist. Do not use for a specific research task's domain science unless Praxist system context is needed first.
Source: `skills/praxist-onboarding/SKILL.md` |
| `praxist-runtime-install` | `$praxist-runtime-install` | Install Praxist runtime dependencies and configure user-level Praxist provider credentials for a source checkout or pip-installed environment. Use when the user asks an agent to install or repair Praxist requirements, prepare a Praxist host, create the Praxist Python environment, install the `praxist` CLI, install Claude SDK or official Codex SDK runtime extras, install source-checkout test/dev dependencies, persist API keys or provider settings, verify imports, or diagnose missing dependencies. For source checkouts include the repository test/dev dependency group; for pip package installs keep the install runtime-only. Do not use for docs, task-specific training, dataset, benchmark, or experiment dependencies.
Source: `skills/praxist-runtime-install/SKILL.md` |
| `praxist-scientific-research` | `$praxist-scientific-research` | Gather task-agnostic scientific research context for Praxist task projects using no-key public literature/database/open-access lookup, agent-host web search when available, local project documents, and source/provenance notes. Use when an agent needs to identify domain metrics, benchmarks, prior art, scientific databases, open-access provenance, high-value research directions, or literature-backed hypotheses for a Praxist task without starting a run, changing Praxist core logic, or treating literature as measured task performance.
Source: `skills/praxist-scientific-research/SKILL.md` |
| `praxist-takeover` | `$praxist-takeover` | Orchestrate first-use Praxist onboarding, current task initialization, validation, and detached run launch from a supported agent interface. Use when a user wants a one-command or one-conversation path from an existing runnable research project to a started Praxist run, asks to onboard plus initialize plus start, asks for repo-to-task-to-run setup, or wants the agent to prepare and launch Praxist with minimal interaction. This skill composes onboarding, full task initialization or task-harness repair, runtime checks, control start, and confirmation gates without changing Praxist core.
Source: `skills/praxist-takeover/SKILL.md` |
| `praxist-takeover-codex` | `$praxist-takeover-codex` | Onboard, initialize or repair, validate, and launch a Praxist research task in Codex-native mode through the official Codex SDK runtime and the operator's existing saved ChatGPT login, with catalog-verified gpt-5.6-luna as the default model and without requesting, storing, or using an API key. Use when a user wants a low-interaction no-key Praxist takeover, explicitly requests Codex-native mode, has no provider key, or invokes this skill with no additional text after already logging in to Codex.
Source: `skills/praxist-takeover-codex/SKILL.md` |
| `praxist-task-initialization` | `$praxist-task-initialization` | Convert an existing runnable computer-based research project into a formal Praxist task project, or repair a task harness that fails task-init validation. Use when a user wants an agent to transform AI algorithm, robotics, control, simulation, SLAM, LLM, optimization, or other executable research code into a Praxist task directory with task.yaml, baseline harness, evaluator, baseline performance records, robust metric/ranking policy, protocol-integrity checks, reachable task-justified durable/Pareto retention lanes, role prompts, audit rules, dataset/simulator metadata, high-value research directions, initial-generation DIG plus independently controlled QD, continuous-evolution/Gems research-loop settings, run-report tooling, and hardware-aware or user-selected fixed Praxist run parameters. The skill requires a project that already runs on the current machine or in an available environment/container. Abort when required code, data/simulator assets, or declared runtime dependencies are missing.
Source: `skills/praxist-task-initialization/SKILL.md` |
| `terminal-line-plot` | `$terminal-line-plot` | Draw readable ASCII line charts directly in the terminal from numeric series, command output, CSV, JSON, or manually extracted points. Use when the user asks for a curve, trend line, score progression, leaderboard trend, metric history, or any plot that should be visible in a CLI/chat transcript without opening a GUI or writing image files.
Source: `skills/terminal-line-plot/SKILL.md` |
---
# Core API Reference
This page is generated from public docstrings in `praxist.core`.
## Registry
::: praxist.core.registry
## Task Projects
::: praxist.core.task_project
## Protocol
::: praxist.core.protocol
## Runtimes
::: praxist.core.runtimes
## Modeling
::: praxist.core.modeling
## Budget
::: praxist.core.budget
## Credentials
::: praxist.core.credentials
## Prompt Layout
::: praxist.core.prompt_layout
## Tool Servers
::: praxist.core.tool_servers
## Role Skills
::: praxist.core.role_skills
## Workflow
::: praxist.core.workflow
## Storage
::: praxist.core.storage
## Trajectory
::: praxist.core.trajectory
## Replay
::: praxist.core.replay
## Execution Guards
::: praxist.core.execution_guards
## Source Snapshot
::: praxist.core.source_snapshot
---
# Plugin API Reference
This page documents executable generic plugin boundaries.
## Agent Runtimes
::: praxist.plugins.agent_runtimes.claude_sdk.adapter
::: praxist.plugins.agent_runtimes.codex_sdk.adapter
## API Providers
::: praxist.plugins.model_providers.openrouter.adapter
::: praxist.plugins.model_providers.anthropic_messages.adapter
::: praxist.plugins.model_providers.openai_compatible.adapter
::: praxist.plugins.model_providers.deepseek_alias.adapter
## Workflow Stage
::: praxist.plugins.workflow_stages.research_loop.startup
::: praxist.plugins.workflow_stages.research_loop.stage
::: praxist.plugins.workflow_stages.research_loop.c5_materializer
::: praxist.plugins.workflow_stages.reviewer_stub.adapter
## Tools
::: praxist.plugins.tools.evaluation_tools.adapter
::: praxist.plugins.tools.finding_graph_query.adapter
::: praxist.plugins.tools.frontier_tools.adapter
::: praxist.plugins.tools.memory_tools.adapter
::: praxist.plugins.tools.prior_work_tools.adapter
::: praxist.plugins.tools.literature_lookup.adapter
## Graph Maintainer
::: praxist.plugins.graph_maintainers.finding_graph_mvp.adapter
::: praxist.plugins.graph_maintainers.finding_graph_mvp.engine
::: praxist.plugins.graph_maintainers.finding_graph_mvp.cli
## Budget Policy
::: praxist.plugins.budget_policies.default_basic.policy
---
# CLI and Operator API Reference
This page documents Python operator entrypoints and support scripts.
## Run CLI
::: praxist.run
## Deliverables
::: praxist.deliver
::: scripts.deliver_auto_research
## Task Spec Compatibility
::: praxist.task_spec
## Fake Workflow Fixture
::: praxist.testing.fake_workflow_fixture
---
# Task Template API Reference
Tracked templates demonstrate the task-project layout. They are not a
production task catalog or complete worked examples. Use
[Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates) to choose the
right asset, and inspect [Rocket Booster Recovery](https://praxist.sapient.inc/en/docs/examples/rocket-booster-recovery)
for a complete runnable integration.
## Template Runner
::: templates.tasks.template.runner
## Toy Math Runner
::: templates.tasks.toy_math.runner
## SAM Optimizer Reference Evaluation
::: templates.tasks.sam_optimizer.evaluations.pareto_tiered.evaluator
## SAM Optimizer Reference Evaluation Runner
::: templates.tasks.sam_optimizer.evaluations.pareto_tiered.run
## Template Evaluation
::: templates.tasks.template.evaluations.pareto_tiered.evaluator
## Template Evaluation Runner
::: templates.tasks.template.evaluations.pareto_tiered.run
## Toy Math Evaluation
::: templates.tasks.toy_math.evaluations.pareto_tiered.evaluator
## Toy Math Evaluation Runner
::: templates.tasks.toy_math.evaluations.pareto_tiered.run
## Configuration Source
The executable API objects above are the purpose of this reference page.
Configuration, evidence maturity, Deep Innovation Gate (DIG),
Quality-Diversity (QD), Gems, lane routing, artifact ownership, and scheduler
requirements for templates are defined once in
[Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects). The tracked task descriptors under
`templates/tasks/` are validated fixtures of that contract, not a second schema
definition.
---
# Glossary
This glossary gives short navigation definitions. The linked pages own the full
contracts.
| Term | Meaning | Canonical detail |
|---|---|---|
| Task project | External runnable research problem supplied to Praxist | [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects) |
| Template | Replaceable task-project scaffold or deterministic smoke fixture | [Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates) |
| Example | Complete runnable reference project with task-owned code, evidence, and tests | [Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates) |
| Peer | One research agent working within a generation | [Research Loop](https://praxist.sapient.inc/en/docs/guides/research-loop-variant-generation-flow) |
| Generation | A cohort of peer work followed by a research-planning boundary | [Research Loop](https://praxist.sapient.inc/en/docs/guides/research-loop-variant-generation-flow) |
| Finding | Structured report of observed evidence or a reusable research lesson | [Architecture](https://praxist.sapient.inc/en/docs/concepts/architecture) |
| Validation signal | Compact, non-durable evidence retained for validation, repair, or diagnostic follow-up | [Architecture](https://praxist.sapient.inc/en/docs/concepts/architecture) |
| Incubator | Task-defined durable lower-admission library for complete, credible candidates | [Flexibility Controls](https://praxist.sapient.inc/en/docs/guides/research-loop-flexibility-controls) |
| Frontier | Durable task-defined promoted evidence used by planning and reporting | [Architecture](https://praxist.sapient.inc/en/docs/concepts/architecture) |
| Gems | Compact selected research memory used when a task enables periodic reset | [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects) |
| Principal Investigator (PI) | Independent planning agent that proposes next-generation work from committed evidence | [Panel Topology Prompts](https://praxist.sapient.inc/en/docs/concepts/panel_topology_prompts) |
| Chair | Planning agent that compares PI proposals and commits one coherent agenda | [Panel Topology Prompts](https://praxist.sapient.inc/en/docs/concepts/panel_topology_prompts) |
| Deep Innovation Gate (DIG) | Deep-reasoning innovation process that compares candidate mechanisms before implementation | [Deep Innovation Gate](https://praxist.sapient.inc/en/docs/guides/deep-innovation-gate) |
| Quality-Diversity (QD) | Allocation principle that preserves varied strong candidate plans | [Quality-Diversity Allocation](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator) |
| Herfindahl-Hirschman Index (HHI) | Concentration measure used to diagnose whether planned or realized work collapsed into too few categories | [Quality-Diversity Allocation](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator) |
| Task harness | Task-owned evaluator, baseline, prompts, roles, and evidence contracts | [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects) |
| Codex-native mode | Explicit Codex SDK route authenticated by a saved ChatGPT login | [Quickstart](https://praxist.sapient.inc/en/docs/getting-started/quickstart) |
---
# Praxist User Agreement
- **Agreement version:** 2026-08-28
- **Effective date:** 28 August 2026
## Chapter 1: General Provisions
### 1.1 Scope
This Agreement is a complete and legally binding contract between the user
(the **User**) and Sapient Intelligence Pte Ltd (the **Company**), concerning
the Praxist software and related services (the **Services**). The Services may
include the locally installed Praxist software, command-line interfaces,
packaged agent skills, documentation, updates, support, and optional
Company-operated network services made available with a Praxist release.
### 1.2 Acceptance
The User accepts this Agreement by selecting **I agree and continue** in the
Praxist first-use experience, by otherwise recording an explicit acceptance,
or by actually using the Services. The User confirms that they have had an
opportunity to review the complete Agreement, including Appendix A and any
supplementary terms incorporated by reference. A User who does not agree must
cancel setup and stop using the Services.
Praxist stores a minimal acceptance record for the current operating-system
user. It contains the Agreement version and digest, acceptance time, and
whether the choice was made directly or relayed by an Agent. This record stays
on the local machine and is not included in optional product-usage events.
### 1.3 Governing framework
This Agreement is formulated in accordance with the laws of the Republic of
Singapore and other applicable laws and regulations. Its purpose is to define
the rights and obligations of both parties and protect their legitimate rights
and interests.
### 1.4 Revisions
The Company may revise this Agreement in response to business development,
changes in applicable law, or changes to the Services. The current Agreement
will be identified by version and published with the relevant release and in
the official Praxist documentation. Where renewed acceptance is required,
Praxist will present the revised Agreement before continuing first-use setup.
Continued use after the applicable notice or acceptance process constitutes
acceptance of the revised Agreement. A User who refuses a revision must stop
using the affected Services.
## Chapter 2: Eligibility, Installation, and Credential Management
### 2.1 Eligibility
Users must be natural persons with full civil capacity, legal persons, or other
organizations legally capable of entering this Agreement. A minor may use the
Services only with the consent and supervision of a legal guardian, who bears
responsibility for that use to the extent required by law.
### 2.2 Installation and configuration
The current Praxist software does not require a separate Praxist account for
local installation or local research operation. The User is responsible for
providing accurate configuration, selecting an authorized model-provider or
runtime account, and ensuring that the research project and its dependencies
may lawfully be used. If a future Company-operated service requires
registration, the User must provide true, accurate, and complete registration
information and must not impersonate another person, create accounts in bulk,
or otherwise abuse registration.
### 2.3 Third-party accounts and credentials
Accounts, subscriptions, API keys, and saved login credentials used by Praxist
may be issued by third-party providers. Their ownership and permitted use are
governed by the User's agreement with the relevant provider. Praxist stores
provider configuration locally when the User asks it to do so and does not
claim ownership of the User's third-party account. The User shall not transfer,
rent, lend, sell, or misuse any Company-operated account if such an account is
provided separately.
### 2.4 Security
The User shall safeguard passwords, API keys, saved login credentials, and
other authentication material and is responsible for operations performed
with those credentials. The User shall promptly notify the Company and the
relevant third-party provider of suspected unauthorized access. The Company
may provide reasonable assistance but is not liable for loss caused by the
User's failure to protect credentials, except where applicable law provides
otherwise.
## Chapter 3: Service Content and Usage Code of Conduct
### 3.1 Service content
Praxist provides a task-agnostic framework for measurable,
computer-executable research. It can coordinate agent runtimes, task-owned
experiments, evidence, research planning, and reporting. Unless expressly
stated otherwise, Praxist does not supply the User's research project,
datasets, simulators, evaluator, task-specific dependencies, scientific
acceptance criteria, computing resources, or third-party model service. The
specific Services available to a User are those included in the installed
release or displayed in official documentation. The Company may add, remove,
or improve Services and will publish material changes through official release
materials or documentation.
### 3.2 Usage rules
The User shall comply with applicable law, public order, and good morals and
shall not use the Services to:
1. create or disseminate content prohibited by applicable law, including
unlawful political propaganda, threats to national security, illegal
gambling, malicious hacking tools, obscene content, or unlawful
discriminatory content;
2. infringe intellectual property, portrait, reputation, privacy, or other
legitimate rights, impersonate another person, or disclose another person's
private information without authority;
3. commit fraud, extortion, harassment, deception, or other illegal acts;
4. maliciously attack, crack, disrupt, tamper with, or steal data from Praxist
or any connected service;
5. use a Service outside the scope authorized by the Company or a relevant
third-party provider; or
6. violate applicable national or regional laws, regulations, sanctions, or
published usage rules.
### 3.3 Service restrictions
The Company may limit usage of Company-operated online Services based on
service capacity, security, legal requirements, or abusive use. Third-party
model providers and infrastructure providers may impose their own quotas and
usage limits. This clause does not give the Company remote control over the
User's lawful local computing resources or task project. The Company may
suspend or terminate access to a Company-operated Service for material breach
without compensation except where applicable law requires otherwise.
## Chapter 4: Intellectual Property Rights
### 4.1 Praxist materials
The Company retains rights in Company-authored software, trademarks, patents,
algorithms, service designs, interface designs, documentation, and written
content to the fullest extent permitted by law. Use and redistribution of
software or documentation are also subject to the license terms accompanying
the relevant distribution. Third-party software, models, datasets, and other
materials remain subject to their respective owners' rights and licenses.
Nothing in this Agreement overrides an applicable open-source or third-party
license.
### 4.2 User-generated content
Unless the parties enter a separate written agreement, the User retains the
copyright and other rights they hold in research inputs, task projects, and
content generated through the Services. The User remains responsible for any
third-party terms that apply to a model, dataset, runtime, or other component
used to produce that content.
### 4.3 Local inputs and voluntary submissions
Local processing by Praxist does not transfer the User's rights in task
content, prompts, research results, files, or project paths to the Company.
Those materials are not included in optional product-usage events described in
Appendix A. If the User separately and voluntarily submits material to the
Company for support, feedback, or another requested service, the User grants a
non-exclusive license limited to providing that service and improving Praxist,
subject to applicable privacy duties and any separate written terms.
### 4.4 User warranty
The User warrants that submitted or processed content is within the User's
right to use and does not infringe third-party rights. The User shall bear
liability for disputes and losses caused by content the User had no right to
use, including legally recoverable damages, litigation costs, and attorney
fees.
## Chapter 5: Suspension, Termination, and Modification of Services
### 5.1 Availability
Company-operated online Services may be suspended for maintenance, upgrades,
malfunctions, force majeure, security incidents, or legal requirements. The
Company will provide reasonable notice where practicable and will announce
restoration when appropriate. Locally installed Praxist software may remain
available, but its operation can depend on third-party runtimes, providers,
networks, or infrastructure outside the Company's control.
### 5.2 User breach
If the User materially breaches this Agreement, uses a Company-operated
Service unlawfully, or provides false information where registration is
required, the Company may restrict or terminate access to that Service and any
associated Company-operated account. This does not authorize the Company to
delete the User's local research project. Local and hosted data, if any, will
be handled under the applicable documentation and law.
### 5.3 Service changes
The Company may modify or terminate part or all of the Services as business or
legal requirements change. Material changes will be published through official
release materials, documentation, or service notices and take effect after any
required notice period. Where a paid Company-operated Service is terminated,
the Company will handle outstanding matters according to the applicable order
terms and law.
## Chapter 6: Rights and Obligations of Both Parties
### 6.1 Company rights and obligations
The Company may:
1. provide the Services under this Agreement, manage Company-operated
Services, and respond to misuse;
2. revise this Agreement and official rules in accordance with Section 1.4;
3. use reasonable efforts to maintain Company-operated systems, publish
software updates, and respond to reasonable feedback;
4. protect personal information in accordance with applicable law and avoid
unauthorized disclosure or misuse;
5. refrain from using the Services to conduct illegal activity or infringe the
legitimate rights of Users or third parties; and
6. collect the bounded product-usage data in Appendix A only after a separate,
explicit opt-in. Acceptance of this User Agreement alone does not enable
product-usage collection.
### 6.2 User rights and obligations
The User may and shall:
1. use the Services under this Agreement and submit reasonable suggestions;
2. exercise applicable rights to inquire about, correct, or delete personal
information and to close any Company-operated account;
3. comply with this Agreement and refrain from illegal acts or conduct that
harms the Company or third parties;
4. safeguard credentials and accept responsibility for authorized operations
performed with them;
5. report material faults or vulnerabilities responsibly and refrain from
malicious exploitation; and
6. comply with relevant terms when using third-party providers or services
through Praxist.
## Chapter 7: Liability for Breach of Contract
### 7.1 User breach
If the User breaches this Agreement, the Company may suspend or restrict
Company-operated Services or terminate a Company-operated account. The User
shall compensate the Company for losses recoverable under applicable law,
including direct losses and reasonable litigation and attorney fees.
### 7.2 Company breach
If the Company breaches this Agreement by failing to perform an applicable
service obligation or infringing the User's legitimate rights, the Company
shall bear liability required by law. To the extent permitted by law, the
Company is not liable for indirect losses or lost anticipated profits, and its
aggregate liability shall not exceed the service fees the User actually paid
to the Company for the affected Service.
### 7.3 Force majeure
Neither party is liable for a failure caused by force majeure, including
earthquakes, floods, typhoons, war, policy changes, widespread system failure,
or cyberattack, to the extent recognized by law. The affected party shall give
prompt notice where practicable and take reasonable steps to reduce loss.
## Chapter 8: Dispute Resolution
### 8.1 Governing law
The formation, performance, interpretation, and dispute resolution of this
Agreement are governed by the laws of Singapore.
### 8.2 Arbitration
The parties shall first attempt to resolve any dispute arising out of or in
connection with this Agreement through good-faith negotiation. If negotiation
fails, either party may submit the dispute to the Singapore International
Arbitration Centre (SIAC) for arbitration under its rules then in force.
## Chapter 9: Miscellaneous Provisions
### 9.1 Severability
If any clause is held invalid or unenforceable, the remaining clauses remain in
force to the extent permitted by law.
### 9.2 Supplementary terms
Matters not covered here may be governed by a separate written supplementary
agreement, which will have the same legal effect when validly entered into by
the parties.
### 9.3 Incorporated documents
Appendix A, applicable service-specific terms, and official notices expressly
incorporated into this Agreement form part of it. Product documentation that
only explains operation does not silently expand the data-collection scope in
Appendix A.
### 9.4 Term and interpretation
This Agreement takes effect when the User records acceptance or begins using
the Services and remains effective until use and any Company-operated account
or Service relationship end, subject to clauses that survive by their nature.
The Company may interpret and revise this Agreement subject to applicable law
and Section 1.4.
- **Company / service operator:** Sapient Intelligence Pte Ltd
- **Contact:** praxist@sapient.inc
- **Release date:** 28 August 2026
Appendix A is the separately authored [Praxist User Data Collection
Notice](https://praxist.sapient.inc/en/docs/legal/product-usage-data-notice). It is incorporated into this Agreement,
but product-usage collection remains disabled unless the User gives the
separate opt-in described there.
---
# Praxist Privacy Notice (Product Usage Data)
**Last updated:** 27 August 2026
**Related documents:** [Fair Source License Agreement (Version 1.0)](https://github.com/sapientinc/praxist/blob/main/LICENSE.md);
[Praxist User Data Collection Notice](https://praxist.sapient.inc/en/docs/legal/product-usage-data-notice) (Notice
version 3); [Product Usage Technical Documentation](https://praxist.sapient.inc/en/docs/operations/DOCUMENTATION)
---
## 1. Overview
This Notice applies only to Praxist's **optional product-usage data collection**. The data controller is Sapient Intelligence Pte Ltd ("we", "us").
Core principles:
- **Voluntary.** Collection is predicated on your explicit consent and is never mandatory.
- **Revocable.** You may withdraw consent at any time; collection stops immediately upon withdrawal.
- **No impact on use.** Declining or withdrawing consent does not affect the installation of Praxist or any research functionality.
- **Pseudonymized and minimized.** We collect only pseudonymized lifecycle data within a closed schema — never any task content, input data, or research results.
## 2. Preconditions for Collection (Voluntary Consent Mechanism)
Praxist collects product-usage data only when **all three** of the following conditions are met:
1. the installed build contains an approved collection transport;
2. product-usage collection is enabled for that build; and
3. you separately and explicitly select **"Share product usage"** after this Notice is made available for review, or reply `Yes` / `Agree` to consent prompt shown during installation.
Please note:
- Accepting the Praxist User Agreement or the Fair Source License Agreement does **not** constitute consent to data collection; the two are independent;
- when no choice has been made (status "unset"), **nothing is collected or uploaded**;
- in non-interactive environments (no terminal input), no consent prompt is shown and collection remains off;
- the same rule applies to Agent-assisted installation: only `Yes` / `Agree` (consent) and `No` / `Disagree` (refusal) are recognized; any other wording is treated as no consent given;
- this Notice is versioned (currently Notice version 3). When the Notice changes, its version number increases; consent you previously recorded does not carry over to a new version and will be requested again.
## 3. What We May Collect (If You Consent)
The product-usage protocol uses a **closed schema** (Schema V2): only the fields below are permitted, and any additional field is rejected by both the client and the server.
### 3.1 Common lifecycle fields
| Field | Description |
| --- | --- |
| `schema_version` | Version of the product-usage event structure (currently 2) |
| `praxist_version` | Public Praxist version |
| `consent_notice_version` | The Notice version you consented to |
| `environment_id` | Environment identifier: a random UUID generated locally, stable across Research Runs within one Praxist environment (see Section 5) |
| `telemetry_run_id` | Independent random identifier for a single Research Run |
| `event_id` | Random identifier for a single event, used for correlation and deduplication |
| `event_sequence` | Sequence number within the same Research Run |
| `event_type` | Lifecycle event type (see Section 3.2) |
| `occurred_at` | Client-side event time (UTC, to the second) |
| `error_summaries` / `error_summaries_truncated` | Bounded structured error-category counts, and whether the bounded list was truncated (see Section 3.3) |
### 3.2 The four lifecycle events
| Event | What it records |
| --- | --- |
| `run_started` | The generation ordinal, the planned Peer count, and aggregate counts of Peers in the planning, running, completed, cancelled, failed, and unknown states at the run-start boundary |
| `generation_finished` | The same fields recorded at a durable generation boundary (state counts sum to the planned Peer count; "completed" means only that a Peer lifecycle returned normally — it does not assert that any scientific result is valid) |
| `run_finished` | Active run duration in complete minutes (capped at 43,200 minutes, i.e. 30 days) and whether the cap was reached |
| `run_reconciled` | The same duration fields, recorded when a previously unfinished run is resumed and trustworthy terminal processing is later completed |
### 3.3 Structured error summaries
Each lifecycle event may carry at most 16 grouped error summaries. Each group may contain only the following closed values:
- `scope`: run / generation / peer
- `stage`: setup / launch / execution / finalization / reconciliation
- `error_type`: configuration / resource / orchestration / runtime / external_dependency / storage / unknown
- `error_code`: PRX-CAPACITY / PRX-PEER-LAUNCH / PRX-PEER-RUNTIME / PRX-RUNTIME / PRX-RUN-FAILED / PRX-UNKNOWN
- `reason_code`: auth_error / quota_exhausted / rate_limited / timeout / provider_unavailable / runtime_error / tool_unavailable / invalid_request / budget_denied / budget_expired / capacity_unavailable / process_start_failed / state_unreadable / unexpected_termination / unknown
- `count`: number of matching errors (capped at 65,535), plus whether the cap was reached
This structure is **technically incapable** of containing raw error messages, logs, stack traces, provider responses, or arbitrary text.
### 3.4 Time fields
- `occurred_at` is generated by the local Praxist client at a lifecycle milestone, converted to UTC using the local system clock (and may therefore reflect clock inaccuracy);
- `received_at` is added by the server-side Collector after validation and cannot be supplied or altered by the client. It represents arrival time and is used only for storage management and retention calculation.
## 4. What We Do Not Collect
Product-usage events do **not** include:
- research task content, prompts, research results, files, filenames, project paths, or commands;
- environment variables, API keys, saved login credentials, logs, stack traces, raw error messages, or arbitrary text;
- model names, service-provider names, provider responses, or account information;
- names, email addresses, operating-system details, hardware information, Python version, client time zone, cookies, or arbitrary request headers;
- individual Peer identities or individual Peer outputs;
- IP addresses (event bodies never contain them; the temporary handling of network connection information is described in Section 6).
## 5. Pseudonymization Methodology
Praxist uses **pseudonymization**, not full anonymization:
- **Generation.** `environment_id`, `telemetry_run_id`, and `event_id` are all randomly generated via standard UUIDv4 (`uuid4()`);
- **No derivation.** These identifiers are not derived from usernames, accounts, device serial numbers, MAC addresses, IP addresses, hostnames, project paths, or task content, and they are not simple incrementing sequences;
- **Persistence.** `environment_id` is generated the first time an environment needs one and stored locally in `environment.json` (readable and writable only by the current user, permission 0600); it remains stable across runs within that environment. `telemetry_run_id` and `event_id` are generated per run and per event;
- **Local path hashing.** Local run-state files are named with the SHA-256 hash of the run path; these files **stay on your machine and are never uploaded**;
- **Honest characterization.** Because `environment_id` is stable across runs, events from the same environment could in theory be linked to one another — this is pseudonymized data, not fully anonymous data. However, it cannot directly identify you or your device, and we commit to **never using it in any way** to identify an individual (consistent with our commitment in Section 7).
## 6. Network and Transmission
- **Production endpoint:** `https://telemetry.theaiscientist.com/v1/events` (HTTPS encryption, server certificate verification, no redirect following);
- Requests carry only a fixed, protocol-level User-Agent (`Praxist-Product-Usage/2`); each request is capped at 32 KB and at most 50 events per batch; network timeout is 2 seconds;
- Transmission failures never block or affect Research Runs; events that fail to send are stored locally and delivered later when the network is available (see Section 8);
- **Server-side handling of connection information.** Product-usage event bodies contain no IP addresses, cookies, or arbitrary request headers. Network services necessarily process connection information transiently to deliver requests, protect the service, and apply rate limits. The Collector deployment disables access logging and strips forwarded IP, cookie, and client User-Agent headers before application processing; none of them are persisted as product-usage event data.
## 7. Purposes of Use
Collected product-usage data is used **only** for:
1. analyzing product reliability;
2. improving Praxist's features, performance, and user experience; and
3. anonymous, aggregate-level statistics that are never presented in a way that could identify a particular user.
We commit that we will **not**:
- sell product-usage data to any third party;
- use it for advertising or marketing; or
- link it with any other dataset about you or your users in order to identify an individual.
## 8. Storage, Retention, and Deletion
- **Server-side storage.** Managed PostgreSQL, accessed over a private endpoint; no public internet database exposure;
- **Retention.** Delivered raw events are retained for **at most 180 days** from server receipt and are then deleted by a scheduled retention process (the job runs at least once daily; by design, events enter the deletion window on day 179, leaving one day of scheduling slack so the stated 180-day ceiling is never exceeded);
- **Local unsent events.** Stored in the local SQLite outbound queue (`outbox.sqlite3`, readable and writable only by the current user), cleared after successful delivery, and deleted immediately upon withdrawal;
- **Honest note.** The server does not offer an interface to delete **already-delivered** events by environment identifier. Withdrawal stops future collection and deletes local unsent events; delivered events are deleted automatically when the retention period expires.
## 9. Your Rights and How to Exercise Them
| Action | Command |
| --- | --- |
| View the complete Notice text | `praxist product-usage notice` |
| Check current consent status | `praxist product-usage status --json` |
| Record consent | `praxist product-usage consent` (or select "Share product usage" during first use) |
| **Withdraw consent at any time** | `praxist product-usage withdraw` |
- Withdrawal immediately stops all future capture and deletes all local unsent events;
- Declining or withdrawing **does not affect** the installation, operation, or any research functionality of Praxist;
- You may also exercise your rights of access, rectification, deletion, or complaint through the contact point in Section 13.
## 10. Data Recipient and Cross-Border Arrangements
- **Data recipient:** Sapient Intelligence Pte Ltd (the same legal entity as the Licensor under the license agreement)
- **Server location:** Johor, Malaysia
- **Cross-border transfer:** If you are located in mainland China and choose to opt in, your consent constitutes authorization for the transfer of the above data — which contains no names, contact details, accounts, IP addresses, or content data (see the closed lists in Sections 3 and 4) — to the Collector in Malaysia, limited to the fields enumerated in the Section 3 closed schema. Users in the EU/EEA are covered by Section 10.1.
### 10.1 Supplementary Notice for Users in the European Union and EEA (GDPR)
Praxist is distributed globally, and users in the EU/EEA may likewise choose to opt in to product-usage collection. Under the GDPR, pseudonymized data remains personal data, and we apply the following rules to such data:
- **Legal basis.** Your explicit consent only (GDPR Art. 6(1)(a)) — data collection is off by default and is enabled only when you actively select "Share product usage";
- **Withdrawal.** You may withdraw consent at any time via `praxist product-usage withdraw`; withdrawal does not affect the lawfulness of processing carried out on the basis of consent before its withdrawal (Art. 9(3));
- **Cross-border transfer.** Data is transferred to and processed by the Collector in Johor, Malaysia. This transfer relies on the explicit consent you give at opt-in (the derogation in GDPR Art. 49(1)(a)) and is limited to the fields enumerated in the Section 3 closed schema of this Notice;
- **Your rights.** Access, rectification, restriction of processing, data portability, objection, and erasure. Send access, rectification, or deletion requests to praxist@sapient.inc;
- **Honest note on erasure.** As described in Section 8, the server currently is not able to search or delete **already-delivered** events by environment identifier; delivered events are deleted automatically at most 180 days after receipt. You may include the `environment_id` from your local `environment.json` in a deletion request to verify ownership of the environment, and we will manually process your request after verification;
- **We do not:** sell personal data, use it for advertising or marketing, or carry out automated decision-making with legal or similarly significant effects;
- **Supervisory authority.** You have the right to lodge a complaint with the data protection supervisory authority of your member state.
## 11. Data Security
- Production traffic is HTTPS-encrypted end to end with certificate verification, and redirects are refused;
- Closed schema: both the client and the server reject any field outside the schema;
- The server enforces global and per-client rate limits and a storage capacity ceiling; a master ingestion switch can shut down the collection entry entirely in an emergency;
- Local consent records, the environment identifier, and the outbound queue are readable and writable only by the current user (0600/0700 permissions);
- Events are deleted automatically when the retention period expires — no action required from you;
- Security contact: praxist@sapient.inc
## 12. Changes to This Notice
We may revise this Notice and the in-product notice as the product evolves. When the notice content changes, the Notice version number increases; consent you recorded is valid only for the version it was given against, and renewed consent will be requested at first use after an upgrade. If this Notice and the in-product notice diverge, the version presented to you at the time of consent prevails.
## 13. Contact Us
- Privacy contact: praxist@sapient.inc
- Company: Sapient Intelligence Pte Ltd (incorporated in Singapore)
- Commercial licensing and license-agreement matters: see the contact details at the end of the Fair Source License Agreement (Version 1.0) (praxist@sapient.inc)
---
# Appendix A: Praxist User Data Collection Notice
**Notice version:** 3
**Effective date:** 19 August 2026
## 1. Separate and Optional Consent
The User may help improve Praxist by sharing pseudonymized (pseudonymous)
product-usage data. Praxist collects it only when all of the following are
true:
1. the installed build contains an approved collection transport;
2. product-usage collection is enabled for that build; and
3. the User separately selects **Share product usage** or otherwise provides an
explicit supported opt-in after this Notice is made available for review.
Accepting the Praxist Fair Source License and User Agreement does **not**
provide this optional consent.
If the installed build has no approved collection capability, or if consent is
unset or denied, no product-usage events are collected or sent. Research
operation is not reduced when the User declines. A temporary network or
collector outage after opt-in may leave bounded events in the local outbox for
later delivery; it does not affect the Research Run.
Development builds send authorized events to the fixed Praxist development
collector at `http://45.78.201.249/v1/events`. This development transport is
plain HTTP and is intended only for internal development or test data. Formal
releases send authorized events to the Praxist production collector at
`https://telemetry.theaiscientist.com/v1/events` using HTTPS certificate
verification and without following redirects.
## 2. Data Praxist May Collect After Opt-In
The product-usage protocol has a closed schema. It may collect only the fields
described below.
### 2.1 Common lifecycle fields
- `schema_version`: version of the product-usage event structure;
- `praxist_version`: public Praxist version;
- `consent_notice_version`: version of this Notice;
- `environment_id` (Environment ID): locally generated random identifier that
remains stable across Research Runs in one Praxist environment;
- `telemetry_run_id`: separate random identifier for one research run;
- `event_id`: random identifier for correlation and deduplication of one event;
- `event_sequence`: sequence number within the same research run;
- `event_type`: one of the lifecycle events below;
- `occurred_at`: client-side event time in UTC to second precision;
- `error_summaries`: bounded structured error-category counts; and
- `error_summaries_truncated`: whether the bounded error list was truncated.
The random identifiers are not derived from usernames, accounts, device serial
numbers, MAC addresses, IP addresses, hostnames, project paths, or task
content, and are not simple incrementing identifiers.
### 2.2 Lifecycle events
Praxist may emit four event types:
1. **`run_started`** records the generation ordinal, planned Peer count, and
aggregate counts of Peers in planning, running, completed, cancelled,
failed, and unknown states at the run-start boundary.
2. **`generation_finished`** records the same generation and aggregate Peer
state fields at a durable Generation boundary. Peer state counts must sum
to the planned Peer count. `completed` means only that the Peer lifecycle
returned normally; it does not assert that a scientific result is valid.
`unknown` means that a canonical terminal state could not be confirmed.
3. **`run_finished`** records active run duration in complete minutes and
whether it reached the 43,200-minute (30-day) recording cap. Duration may be
null when it cannot be determined.
4. **`run_reconciled`** records the same bounded duration fields when a
previously unfinished run is resumed and trustworthy terminal processing is
later completed.
Individual Peer identifiers, outputs, prompts, research conclusions, and raw
error details are not collected.
### 2.3 Structured error summaries
Each lifecycle event may include up to 16 grouped summaries. Each group may
contain only:
- `scope`: `run`, `generation`, or `peer`;
- `stage`: `setup`, `launch`, `execution`, `finalization`, or
`reconciliation`;
- `error_type`: `configuration`, `resource`, `orchestration`, `runtime`,
`external_dependency`, `storage`, or `unknown`;
- `error_code`: `PRX-CAPACITY`, `PRX-PEER-LAUNCH`, `PRX-PEER-RUNTIME`,
`PRX-RUNTIME`, `PRX-RUN-FAILED`, or `PRX-UNKNOWN`;
- `reason_code`: `auth_error`, `quota_exhausted`, `rate_limited`, `timeout`,
`provider_unavailable`, `runtime_error`, `tool_unavailable`,
`invalid_request`, `budget_denied`, `budget_expired`,
`capacity_unavailable`, `process_start_failed`, `state_unreadable`,
`unexpected_termination`, or `unknown`;
- `count`: count of matching errors, capped at 65,535; and
- `count_capped`: whether the count reached that cap.
The structure cannot contain raw error messages, logs, stack traces, provider
responses, or arbitrary text.
### 2.4 Time fields
`occurred_at` is generated by the local Praxist client at a lifecycle
milestone, converted to UTC using the local system clock, and may reflect clock
inaccuracy. The Collector adds `received_at` after validating an event. This
server receipt time cannot be supplied or changed by the client. It represents
arrival time, not task completion time, and is used for storage management and
retention rather than local duration calculation.
## 3. Data Praxist Does Not Collect
Product-usage events do not include:
- research task content, prompts, research results, files, filenames, project
paths, or commands;
- environment variables, API keys, saved login credentials, logs, stack
traces, raw error messages, or arbitrary text;
- model names, service-provider names, provider responses, or account
information;
- names, email addresses, operating-system details, hardware information,
Python version, client time zone, cookies, or arbitrary request headers; or
- individual Peer identities or individual Peer outputs.
## 4. Network Information
Product-usage event bodies do not include IP addresses, cookies, or arbitrary
request headers. Network services necessarily process connection information
temporarily to deliver a request, protect the service, and apply rate limits.
The Praxist Collector deployment disables access logging, strips forwarded IP,
cookie, and client User-Agent headers before application processing, and does
not persist them as product-usage event data. The client sends only a fixed,
protocol-level User-Agent header.
## 5. Retention and Withdrawal
Delivered raw events are retained for no more than 180 days and are then
deleted by the scheduled retention process. The User may run:
```bash
praxist product-usage withdraw
```
Withdrawal immediately disables future capture for the current user and
deletes unsent local events. It does not delete already delivered events;
those remain only until the scheduled retention period expires. The current
state is available through:
```bash
praxist product-usage status --json
```
During Agent-assisted OOBE, the supported explicit sharing replies are `Yes`
and `Agree`; the supported refusal replies are `No` and `Disagree`. An Agent
must not infer a choice from other language.
---
# Documentation Policy
The versioned Markdown under `docs/` is the only authored product
documentation source. The static website, generated reference pages, search
index, `llms.txt`, and `llms-full.txt` are derived from it. Praxist does not use
a separately edited GitHub Wiki.
## One Fact, One Owner
| Information | Sole owner | Other pages may |
|---|---|---|
| CLI arguments and defaults | `praxist.cli` parser code | Link to generated CLI reference |
| Skill name and activation description | Each `skills/*/SKILL.md` front matter | Link to generated Skills reference |
| Package installation, install commands, and filesystem effects | [Installation](https://praxist.sapient.inc/en/docs/getting-started/installation) | Link; the root landing page may show one canonical command |
| User-facing first-run/OOBE sequence | [Quickstart](https://praxist.sapient.inc/en/docs/getting-started/quickstart) | Name the lane and link |
| Agent-managed OOBE implementation | [Agent OOBE Runbook](https://praxist.sapient.inc/en/docs/agents/oobe-install) | Link without duplicating agent instructions |
| Project prerequisites, research brief, and takeover stages | [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task) | Show one takeover invocation and link |
| Goal-to-skill map and skill installation locations | [Agent Skills](https://praxist.sapient.inc/en/docs/user-guide/skills) | Name a relevant skill and link |
| Task schema, precedence, and scientific ownership | [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects) | Explain how a mechanism consumes the task contract |
| Template/example distinction and example materialization commands | [Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates) | Identify an asset and link without redefining the boundary |
| Rocket Booster Recovery scientific details | `examples/rocket_booster_recovery/README.md` | Explain discovery and launch without duplicating its protocol |
| Rocket Booster Recovery (Rust) scientific details | `examples/rocket_booster_recovery_rust/README.md` | Explain discovery and launch without duplicating its protocol |
| Lifecycle semantics | [Direct CLI Operations](https://praxist.sapient.inc/en/docs/guides/operators) | Show one quickstart command and link |
| Core/plugin/task boundary and artifact roles | [Architecture](https://praxist.sapient.inc/en/docs/concepts/architecture) | Summarize purpose and link |
| Configuration ingress and precedence | [Configuration Discipline](https://praxist.sapient.inc/en/docs/concepts/config_discipline) | State a local input and link |
| Agent-session and prompt-layout mental model | [Runtime Model](https://praxist.sapient.inc/en/docs/concepts/runtime-model) | Explain a mechanism-specific effect and link |
| Runtime adapter capabilities | [Agent Runtimes](https://praxist.sapient.inc/en/docs/guides/agent-runtimes) | Identify a selected runtime and link |
| API provider shapes | [API Providers](https://praxist.sapient.inc/en/docs/guides/model-providers) | Identify a selected provider and link |
| Open-source model API shortlist and selection criteria | [Open-Source Model APIs](https://praxist.sapient.inc/en/docs/guides/open-source-model-apis) | Link without duplicating the shortlist |
| Authentication and credential precedence | [Credentials](https://praxist.sapient.inc/en/docs/guides/credentials) | State a prerequisite and link |
| Research-loop sequence | [Research Loop](https://praxist.sapient.inc/en/docs/guides/research-loop-variant-generation-flow) | Refer to a stage without redefining the sequence |
| Maturity, close, incubator, peer-mix, and launch-freeze behavior | [Flexibility Controls](https://praxist.sapient.inc/en/docs/guides/research-loop-flexibility-controls) | Show task configuration only where Task Projects owns the combined profile |
| Deep Innovation Gate (DIG) behavior and artifacts | [Deep Innovation Gate](https://praxist.sapient.inc/en/docs/guides/deep-innovation-gate) | State whether DIG is active and link |
| Quality-Diversity (QD) allocation | [Quality-Diversity Allocation](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator) | Refer to the selected path and link |
| Peer-session memory | [Peer Memory](https://praxist.sapient.inc/en/docs/guides/peer-local-structured-memory-long-context) | Refer to memory as context and link |
| Experiment admission and resource ownership | [Central Experiment Scheduler](https://praxist.sapient.inc/en/docs/guides/central-resource-scheduler) | State the selected profile and link |
| Tool catalog and frontier-tool behavior | [Tool Servers](https://praxist.sapient.inc/en/docs/guides/tool-servers) | Name a tool and link |
| Literature source/provenance policy | [Scientific Literature Lookup](https://praxist.sapient.inc/en/docs/guides/scientific-literature-lookup) | State whether lookup is enabled and link |
| Human-readable report triggers and semantics | [Run Reports](https://praxist.sapient.inc/en/docs/guides/user-facing-reports-and-init) | Link to a generated report or this guide |
| Usage measurement and formulas | [Cost Estimation](https://praxist.sapient.inc/en/docs/guides/costs) | Link to measured artifacts and this guide |
| Lossless token-saving mechanisms | [Cost Optimization](https://praxist.sapient.inc/en/docs/guides/cost-optimization) | State that a route uses the policy and link |
| Contributor contract | `AGENTS.md` | Link without redefining it |
| Software license | Root [`LICENSE.md`](https://github.com/sapientinc/praxist/blob/main/LICENSE.md) | Link without restating license terms |
| User Agreement | [Praxist User Agreement](https://praxist.sapient.inc/en/docs/legal/user-agreement) | Link without restating legal terms |
| Product-usage privacy policy, processing purposes, retention, and user rights | [Privacy Notice](https://praxist.sapient.inc/en/docs/legal/PRIVACY) | Link without restating the policy |
| Exact versioned product-usage consent text | [Praxist User Data Collection Notice](https://praxist.sapient.inc/en/docs/legal/product-usage-data-notice) | The CLI loads this same package resource; other pages identify it and link |
| Product-usage consent commands and collector operation | [Product Usage Controls](https://praxist.sapient.inc/en/docs/operations/product-usage) | Link to the operational procedure |
| Product-usage implementation, endpoint, storage, and audit map | [Product Usage Technical Documentation](https://praxist.sapient.inc/en/docs/operations/DOCUMENTATION) | Link without duplicating implementation details |
| Machine product-usage event schema | `praxist/product_usage/protocol.py` and checked-in JSON Schema | Technical and legal pages explain the contract; code remains authoritative for exact validation |
| Hosted documentation URL | `praxist.cli.docs.DOCUMENTATION_URL` | Mirror it in checked package/site metadata and link to it |
Tutorials sequence actions. Guides explain procedures. Concept pages explain
mental models. Reference pages enumerate machine contracts. The root README is
a product landing page, not a second manual.
A minimal command may appear in a tutorial that needs the action, but its
arguments, defaults, and edge cases remain in the generated reference. A page
may summarize an adjacent contract only far enough to explain its own behavior;
the summary must link to the owner instead of restating the contract.
## Generated Sources
`scripts/build_docs_site.py` creates:
- `docs/reference/cli.md` from the live CLI parser;
- `docs/reference/skills.md` from skill front matter;
- `site/llms.txt` as a compact machine-readable map;
- `site/llms-full.txt` as the complete navigation-ordered corpus.
Generated HTML and LLM exports are not committed. Generated Markdown reference
pages are committed so package users can read them without building the site,
but CI verifies that they exactly match their code-owned inputs.
## Navigation Ownership
Every authored Markdown page must appear exactly once in `mkdocs.yml`. This
makes its primary audience and information role explicit. The docs build
rejects missing pages, duplicate navigation ownership, and broken local links.
## Build
```bash
uv sync --extra docs
uv run python scripts/build_docs_site.py
uv run python scripts/build_docs_site.py --check-generated
```
The build runs in strict mode and does not contact API providers, read API
keys, or start research services.
## Hosted Documentation
Documentation validation runs on every pull request and every push to `main`.
Successful pushes to `main` publish the generated site to the repository's
private GitHub Pages project. Access follows repository read permission and
requires GitHub authentication; the generated HTML remains derived output and
is not committed.
The repository variable `PRAXIST_PAGES_ENABLED=true` is the deployment switch.
Maintainers can also trigger the `docs` workflow manually. Pull requests build
and validate the complete site but never publish it.
The canonical site URL is surfaced through `praxist docs`. Contributors use the
local build commands above only to preview unmerged changes.
---
# Legacy Migration Guide
Legacy migration is allowed only when it preserves behavior while moving code
toward the current core-plugin-task boundaries.
## Migration Pattern
1. Characterize the old behavior with tests.
2. Add the new protocol, plugin, or task boundary.
3. Route execution through the new boundary.
4. Add old-vs-new parity checks.
5. Preserve partial outputs and weak provenance when needed.
6. Delete the old path after the new path is proven.
7. Record migration context in the commit or pull-request description when the
migration changes an architecture contract.
## What Not To Preserve
Do not preserve obsolete package names, duplicate task catalogs, shell-owned
semantics, fake production plugins, or hidden SAM-specific global defaults.
Do not keep compatibility shims after they stop serving an active migration.
## Migration Tests
Use characterization tests for old behavior, unit tests for the new interface,
workflow smoke tests for run shape, and replay tests for artifact consistency.
Long real GPU/API dogfood runs are manual/on-demand gates, not default unit
tests.