# Praxist — Official Product and Documentation Corpus > Praxist is an autonomous research product developed by Sapient Intelligence Pte Ltd. This file is generated from the website's canonical documentation, FAQ, and measured examples. - Product: https://praxist.sapient.inc/en - Company: https://sapient.inc - Source: https://github.com/sapientinc/PRAXIST - Product release date: 28 August 2026 ## Product FAQ ### What is Praxist? Praxist is an autonomous research system for measurable research problems that can be executed on a computer. It turns an already runnable project into a continuous, evidence-driven research run. Across successive generations, parallel research agents develop candidate solutions; evaluators convert results into structured evidence; and a planning panel synthesizes that evidence into the research agenda for the next generation. The cycle continues until the search converges or the budget is exhausted. You provide a runnable project and a measurable objective. Praxist orchestrates the research process that searches for the best-performing solution. ### How is Praxist different from manual tuning or AutoML? AutoML tunes parameters within a predefined search space. Praxist runs the full research loop. Parallel research agents can change methods, architectures, and strategies. Evidence from evaluation shapes the agenda for the next generation, while the Deep Innovation Gate (DIG) and Quality-Diversity (QD) allocation help the system escape local optima. Praxist is closer to a self-directing research team than a search tool. If your researchers are already iterating on a problem manually, Praxist takes over the iteration loop itself. ### Is my project a good fit for Praxist? Praxist delivers the most value when three conditions are met: - **The objective is measurable:** there is at least one metric that meaningfully distinguishes better from worse, with a clear optimization direction. - **The project already runs:** the baseline code, environment, and required data or simulator are in place and work without Praxist. - **The best path forward is unknown.** If a prerequisite is missing, Praxist stops and tells you exactly what is needed. It will not silently download unspecified datasets, invent a simulator, or fabricate baseline performance. That is a deliberate design principle. ### Do I need an API key, and what will it cost? No API key is required in Codex-native mode; Praxist uses your authenticated Codex session. We also recommend using your own API key to access supported model APIs. API costs are set by the provider and vary by model and usage. Total cost also depends on parallelism, the number of generations, and evaluation runtime. For cost-sensitive runs, start with a small representative workload before scaling up. ### How does Praxist protect my code and data? Praxist provides three layers of protection: - **Project isolation:** Praxist does not modify your original project. Run artifacts are stored separately. - **Credentials:** API keys are entered through a masked local prompt and are not exposed in commands, shell history, or conversations. - **Data collection:** Praxist does not collect data used in your experiments. It collects only limited system-level operational information, which you can disable at any time. ### How can I trust that a reported improvement is real? Praxist uses three safeguards: - **Preregistration:** Metrics, evaluation protocols, baselines, and acceptance thresholds are defined before the run. - **Consistent evaluation:** Every candidate is measured through the same evaluator, and invalid or suspicious results are excluded. - **End-to-end provenance:** Every reported improvement includes the evidence and lineage needed to inspect and reproduce it. We recommend reviewing what the selected solution changed and testing it again in your own environment. Praxist's results are designed to be verifiable, and your own validation should be the final test. ### What if Praxist does not improve the result? Praxist does not guarantee a specific metric improvement. It provides a rigorous research process and auditable evidence. If a run does not meet its target, you still receive a negative-result evidence package, an audit report, and recommendations on whether to stop or redirect the research. A negative result can still be valuable: it rules out tested approaches with evidence and helps prevent further investment in an unproductive direction. ### Is Praxist open source, and what terms apply to its outputs? Praxist is licensed under the Fair Source License Agreement 1.0. The precise description is source-available: the complete source code is publicly available and may be viewed, downloaded, and modified. Subject to the license terms, Praxist may be used for internal business purposes and deployed within your own organization. Organizations with aggregate annual revenue, including revenue from affiliates, below US$1 million may use Praxist commercially at no charge. Once annual revenue reaches or exceeds that threshold, the organization must contact the Licensor, Sapient Intelligence Pte Ltd, to negotiate a Commercial License. The revenue threshold does not apply to qualifying teaching and academic research conducted by institutions of higher education, public research institutions, and nonprofit academic research organizations. **Generated outputs:** no attribution is required for internal use. If an output is published externally or otherwise made available to third parties, the product-name attribution "Praxist by Sapient Intelligence" must be retained. This FAQ is a summary only. If it conflicts with the Fair Source License Agreement 1.0, the terms of the license agreement control. ## Measured Examples ### Mobile crop-disease diagnosis - Industry: Agriculture & Environmental Monitoring - Result: ACCURACY 0.8909 → 0.90209 ↑ - Application: Use field photos to identify likely crop disease and route growers toward the right treatment, agronomist, or containment action. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/cassava-leaf-disease-classification/leaderboard.csv ### Classification from images plus engineered measurements - Industry: Agriculture & Environmental Monitoring - Result: LOG LOSS 0.10834 → 0 ↓ - Application: Combine visual evidence with measurements to classify parts, materials, crops, or defects when the labeled dataset is too small for image-only deep learning. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/leaf-classification/leaderboard.csv ### Deploy image models into new regions - Industry: Agriculture & Environmental Monitoring - Result: MACRO F1 0.108 → 0.40915 ↑ - Application: Train on established sites and retain accuracy at new branches, geographies, suppliers, or camera installations where backgrounds and class frequencies shift. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/iwildcam-2019-fgvc6/leaderboard.csv ### Visual identity matching with unknown detection - Industry: Agriculture & Environmental Monitoring - Result: MAP@5 0.32788 → 0.59833 ↑ - Application: Match people, animals, assets, products, or components to known identities from distinctive markings while explicitly handling previously unseen identities. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/whale-categorization-playground/leaderboard.csv ### Next-best-product recommendations - Industry: Commerce, Marketing & Customer Experience - Result: MAP@12 0.02177 → 0.03344 ↑ - Application: Recommend the most relevant products, offers, content, or replenishment items from behavioral history and catalog context. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/h-and-m-personalized-fashion-recommendations/leaderboard.csv ### Identify the words driving customer sentiment - Industry: Commerce, Marketing & Customer Experience - Result: JACCARD 0.71378 → 0.72392 ↑ - Application: Highlight the phrase behind praise, frustration, churn risk, or a complaint so teams see not only the score but the actionable reason. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/tweet-sentiment-extraction/leaderboard.csv ### Predict response or conversion from narrative plus context - Industry: Commerce, Marketing & Customer Experience - Result: AUROC 0.5996 → 0.8414 ↑ - Application: Estimate whether an application, appeal, campaign, lead, donation request, or support message will receive a positive response using both language and metadata. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/random-acts-of-pizza/leaderboard.csv ### Authorship, source, and style attribution - Industry: Content, Media & Communications - Result: LOG LOSS 0.41879 → 0.16327 ↓ - Application: Attribute documents or messages to likely sources, teams, content types, or style profiles for forensics, routing, brand consistency, and provenance analysis. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/spooky-author-identification/leaderboard.csv ### Rich visual attribute tagging - Industry: Content, Media & Communications - Result: MICRO F1 0.627 → 0.68815 ↑ - Application: Generate multiple structured attributes for products, creative assets, real estate, archives, or inspection photos to improve discovery and analytics. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/imet-2020-fgvc7/leaderboard.csv ### Legal Document Transcription - Industry: Education & Public Services - Result: F1 0.5985 → 0.97413 ↑ - Application: Transcribe difficult scanned records into searchable text for legal discovery, court and land-registry archives, document migration, and compliance workflows. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/kuzushiji-recognition/leaderboard.csv ### Consistent scoring of written submissions - Industry: Education & Public Services - Result: QUADRATIC KAPPA 0.82827 → 0.83838 ↑ - Application: Score essays, grant narratives, applications, audits, or quality reviews against a rubric to prioritize human attention and provide faster feedback. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/learning-agency-lab-automated-essay-scoring-2/leaderboard.csv ### Automated retinal screening and referral prioritization - Industry: Healthcare & Life Sciences - Result: QUADRATIC KAPPA 0.88891 → 0.93198 ↑ - Application: Triage screening images by disease severity so clinicians can focus first on patients most likely to need urgent follow-up while preserving human review. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/aptos2019-blindness-detection/leaderboard.csv ### Predict hidden biological or material traits from imaging - Industry: Healthcare & Life Sciences - Result: AUROC 0.52553 → 0.65882 ↑ - Application: Infer an expensive, invasive, or slow-to-measure property from multiple imaging channels, enabling earlier triage and more targeted confirmatory testing. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/leaderboard.csv ### RNA degradation and stability prediction - Industry: Healthcare & Life Sciences - Result: LOG LOSS 0.3631 → 0.22453 ↓ - Application: Predict how a biological or material sequence behaves at every position and under multiple conditions, helping researchers screen designs before costly experiments. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/stanford-covid-vaccine/leaderboard.csv ### Skin cancer detection from dermoscopy images - Industry: Healthcare & Life Sciences - Result: AUROC 0.9128 → 0.94608 ↑ - Application: Prioritize rare, high-risk cases from images and contextual metadata while explicitly optimizing sensitivity under severe class imbalance. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/siim-isic-melanoma-classification/leaderboard.csv ### Molecular and materials property prediction - Industry: Industrial, Energy & Scientific R&D - Result: MEAN-COLUMN RMSLE 0.06988 → 0.04969 ↓ - Application: Screen candidate materials or formulations virtually, reducing the number of simulations and physical experiments needed to find promising designs. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/nomad2018-predict-transparent-conductors/leaderboard.csv ### Remote asset-versus-hazard classification - Industry: Industrial, Energy & Scientific R&D - Result: LOG LOSS 0.20371 → 0.12707 ↓ - Application: Classify ambiguous targets in radar, sonar, thermal, or satellite imagery to protect offshore operations, shipping, borders, and remote infrastructure. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/statoil-iceberg-classifier-challenge/leaderboard.csv ### Real-time 3D hazard and asset detection - Industry: Mobility, Logistics & Geospatial - Result: MAP 0.042 → 0.15199 ↑ - Application: Build a perception layer for vehicles, robots, yards, mines, or warehouses that locates people, equipment, and obstacles in three dimensions so automated systems can act safely. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/3d-object-detection-for-autonomous-vehicles/leaderboard.csv ### Binary visual inspection and routing - Industry: Operations, Risk & Forecasting - Result: LOG LOSS 0.12216 → 0.00098 ↓ - Application: Classify an image into pass/fail, target/non-target, damaged/undamaged, or eligible/ineligible to automate a high-volume first decision. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/dogs-vs-cats-redux-kernels-edition/leaderboard.csv ### Operational case classification and routing - Industry: Operations, Risk & Forecasting - Result: ACCURACY 0.95342 → 0.96295 ↑ - Application: Assign each transaction, account, case, or asset to the right category from structured features for workflow routing and downstream decisions. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/tabular-playground-series-dec-2021/leaderboard.csv ### Recognize work steps and gestures from multimodal sensors - Industry: Operations, Risk & Forecasting - Result: LEVENSHTEIN 0.322 → 0.03374 ↓ - Application: Identify assembly steps, safety gestures, customer interactions, or operator actions by combining video, depth, and audio over time. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/multi-modal-gesture-recognition/leaderboard.csv ### Real-time abuse screening for digital channels - Industry: Security, Trust & Compliance - Result: AUROC 0.77842 → 0.95897 ↑ - Application: Score comments, chats, reviews, and community posts for likely abuse so platforms can warn users, prioritize moderation, or apply policy controls. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/detecting-insults-in-social-commentary/leaderboard.csv ### Policy-aware content moderation - Industry: Security, Trust & Compliance - Result: MEAN-COLUMN AUROC 0.98079 → 0.98803 ↑ - Application: Detect multiple policy violations in comments, reviews, chats, or cases so different behaviors can trigger different workflows rather than one blunt toxic/not-toxic rule. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/jigsaw-toxic-comment-classification-challenge/leaderboard.csv ### Automatic case, document, and ticket tagging - Industry: Software, Knowledge & Productivity - Result: MICRO F1 0.60685 → 0.79589 ↑ - Application: Assign multiple topics to support tickets, knowledge articles, legal documents, or incident reports for routing, search, analytics, and ownership. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/facebook-recruiting-iii-keyword-extraction/leaderboard.csv ### Context-aware text repair and missing-field recovery - Industry: Software, Knowledge & Productivity - Result: LEVENSHTEIN 5.55211 → 5.30562 ↓ - Application: Repair truncated messages, OCR gaps, transcription omissions, or incomplete product and case descriptions using surrounding context. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/billion-word-imputation/leaderboard.csv ### Multilingual customer and employee question answering - Industry: Software, Knowledge & Productivity - Result: WORD JACCARD 0.72756 → 0.88943 ↑ - Application: Answer questions from policies, help centers, product documentation, or case files in lower-resource languages without forcing users into English-only workflows. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/chaii-hindi-and-tamil-question-answering/leaderboard.csv ### Patent phrase similarity and deduplication - Industry: Software, Knowledge & Productivity - Result: PEARSON R 0.851 → 0.87388 ↑ - Application: Find equivalent or related technical concepts across patents, requirements, contracts, product catalogs, and knowledge bases despite different wording. - Evidence source: https://github.com/openai/mle-bench/blob/main/mlebench/competitions/us-patent-phrase-to-phrase-matching/leaderboard.csv ## Technical Documentation # Praxist Praxist coordinates parallel research agents, experiments, evidence retention, and synthesis across generations. It supplies the reusable research process; your task project supplies the science. Read the technical paper, [*Praxist: From Experimental Artifacts to Solution Lineages*](https://arxiv.org/abs/2608.25955), for the system design and evaluation.
![A generation inherits frontier, agenda, and Gems; allocates parallel peers that build artifacts; evaluates those artifacts into typed findings; and synthesizes findings into a frontier update and the next agenda, which the following generation inherits. Artifacts, findings, decisions, and agendas accumulate into a lineage DAG that explains the final artifact.](assets/figures/praxist-overview.svg)

Parallel work becomes measured evidence, durable research state, and a focused agenda for the next generation.

## Research Infrastructure, Explicit Science A coding agent is the recommended interface between the researcher and two deliberately separate systems.
```mermaid flowchart LR HUMAN(("Researcher")) AGENT(["Codex / Claude Code
recommended interface"]) PRAXIST["Praxist
generic research process"] TASK[["Task project
scientific contract"]] HUMAN --> AGENT AGENT --> PRAXIST AGENT --> TASK PRAXIST <--> TASK class HUMAN actor class AGENT interface class PRAXIST system class TASK task ```
| Praxist owns | The task project owns | |---|---| | Peers, generations, orchestration, resource scheduling, and lifecycle | Objective, constraints, baseline, environment, and permitted changes | | Agent runtimes, evidence transport, durable state, replay, and synthesis | Evaluator, metrics, protocol, evidence maturity, prompts, and roles | The boundary keeps Praxist reusable across fields and makes the task project the sole source of domain meaning. [Architecture](https://praxist.sapient.inc/en/docs/concepts/architecture) defines the software boundary; [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects) defines the scientific contract. ## Install Praxist Install and configure Praxist with one command: ```bash python3 -m pip install --index-url https://pypi.org/simple "praxist[agents,codex]" && praxist setup --interactive --install-skills codex ``` Or let an agent install from PyPI and follow the packaged setup runbook: ```text codex --yolo # or: claude --dangerously-skip-permissions Install and configure Praxist using its packaged OOBE runbook. Stop after readiness checks. ``` Installation never selects a project or starts research. Before the separate takeover step, read the [Quickstart](https://praxist.sapient.inc/en/docs/getting-started/quickstart) and [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task). The selected project must already contain the code and local resources needed to run its baseline; Praxist does not invent missing data, simulators, credentials, or measurements. ## Choose A Runtime Profile
- :material-rocket-launch: **Start without an API key** --- Use an existing Codex login through Codex-native mode. [Codex-native profile](https://praxist.sapient.inc/en/docs/getting-started/quickstart#codex-native-mode-no-api-key) - :material-cash-multiple: **Run cost-efficient long research** --- Prefer a high-cache-hit-rate [open-source model API](https://praxist.sapient.inc/en/docs/guides/open-source-model-apis) after a representative quality and cache check. [Model API selection](https://praxist.sapient.inc/en/docs/guides/open-source-model-apis)
## Continue After Setup
- :material-flask-outline: **Prepare a research task** --- Define the research brief, prerequisites, and launch gates. [Your first task](https://praxist.sapient.inc/en/docs/getting-started/first-task) - :material-console: **Operate from the shell** --- Use the direct CLI for lifecycle and monitoring operations. [Direct CLI operations](https://praxist.sapient.inc/en/docs/guides/operators)
## The Research Brief Is the Control Surface Takeover can inspect code and measure an existing baseline, but it cannot infer the researcher's priorities. The brief should identify the objective, credibility standard, allowed resources, exploration policy, and practical run budget. Those decisions shape research direction, experiment throughput, and retention from the first generation onward. [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task#write-the-research-brief) provides a complete example and explains how takeover turns that brief into a validated task project. ## Find a Specific Answer | You want to... | Start here | |---|---| | Install and configure Praxist | [Installation](https://praxist.sapient.inc/en/docs/getting-started/installation) | | Complete setup and hand off a project | [Quickstart](https://praxist.sapient.inc/en/docs/getting-started/quickstart) | | Understand project prerequisites | [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task) | | Use agent workflows | [Agent Skills](https://praxist.sapient.inc/en/docs/user-guide/skills) | | Diagnose a failure or stall | [Troubleshooting](https://praxist.sapient.inc/en/docs/operations/troubleshooting) | | Understand the research loop | [Research Loop](https://praxist.sapient.inc/en/docs/guides/research-loop-variant-generation-flow) | | Configure a task harness | [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects) | | Choose a scaffold or complete reference | [Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates) | | Inspect a complete Python/JAX project | [Rocket Booster Recovery](https://praxist.sapient.inc/en/docs/examples/rocket-booster-recovery) | | Inspect a complete native Rust project | [Rocket Booster Recovery (Rust)](https://praxist.sapient.inc/en/docs/examples/rocket-booster-recovery-rust) | | Extend Praxist | [Developer Guide](https://praxist.sapient.inc/en/docs/guides/contributing) | | Look up an exact command | [CLI Reference](https://praxist.sapient.inc/en/docs/reference/cli) | The [Documentation Policy](https://praxist.sapient.inc/en/docs/about/documentation) identifies the sole owner of each contract. --- # Installation Praxist release CI qualifies Linux on CPython 3.11 and 3.12. The package also targets macOS and other CPython 3.11+ environments, but those combinations are not continuously release-tested. Run `praxist doctor` on every host before launching research. Install Praxist in the Python environment from which Codex or Claude Code will operate. Task-specific packages, datasets, simulators, and accelerator libraries remain in the task environment. ## Prerequisites ```bash python3 --version codex --version # when Codex is the operator interface claude --version # when Claude Code is the operator interface ``` The selected agent interface must already be usable. Authentication requirements for each runtime route are defined in [Credentials](https://praxist.sapient.inc/en/docs/guides/credentials). ## Install And Configure The supported runtime extras install both maintained peer runtimes and the Codex-native integration. Choose the agent application whose bundled skills should be registered; each complete installation command is a single line. ```bash # Codex python3 -m pip install --index-url https://pypi.org/simple "praxist[agents,codex]" && praxist setup --interactive --install-skills codex # Claude Code python3 -m pip install --index-url https://pypi.org/simple "praxist[agents,codex]" && praxist setup --interactive --install-skills claude ``` Run the command in the intended active virtual environment when the host Python is externally managed. `python3 -m pip` keeps the package and `praxist` entrypoint tied to the same interpreter. The base `praxist` package supports inspection and package-level CLI operations; the documented extras are the complete research-runtime installation. The explicit public index prevents an incomplete package mirror from silently omitting a pinned runtime SDK. The command stops if installation, setup, or readiness fails. A successful command also stops at that boundary: it does not select a project or launch research. The [Quickstart](https://praxist.sapient.inc/en/docs/getting-started/quickstart) is the sole description of the OOBE sequence and its agent-managed alternative. ## What Installation Changes The one-line flow: - installs the tested agent-runtime dependencies; - exposes `praxist` in the selected Python environment; - registers bundled skills only for the selected agent host; - writes only the configuration explicitly selected during setup; - materializes writable complete examples outside the package; - runs host diagnostics; and - stops before project selection and research launch. It does not install task training dependencies, CUDA, datasets, simulators, human-facing Codex or Claude Code applications, or a collector service. Codex-native support downloads a platform-specific runtime package of roughly 100-150 MB, so the first pip installation may take several minutes. Same-name operator-owned skills are never replaced silently. Interactive setup offers keep, backup-and-replace, or cancel; non-interactive setup reports the conflict and stops. Open the hosted documentation with: ```bash praxist docs ``` Read the [Quickstart](https://praxist.sapient.inc/en/docs/getting-started/quickstart) and [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task) before using the separate takeover command. ## Writable Complete Examples Read-only package resources are never research workspaces. First-use setup copies bundled complete examples to `${PRAXIST_EXAMPLES_HOME:-~/PraxistExamples}` and prints each writable path. Existing destinations are preserved during upgrades. The discovery and materialization commands, available projects, and asset boundary are owned by [Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates). ## Verify ```bash praxist --version praxist doctor praxist examples list ``` `doctor` checks each detected Praxist-managed skill host. Use `praxist doctor --target codex` or `praxist doctor --target claude` to inspect one host explicitly. If the shell cannot find `praxist`, activate the Python environment used for installation or add that environment's script directory to `PATH`. ## Uninstall Stop active runs, remove Praxist-managed user state, then uninstall the Python package from the same environment: ```bash praxist uninstall --dry-run praxist uninstall python3 -m pip uninstall praxist ``` `praxist uninstall` removes only proven Praxist-managed skills, configuration, local agreement/usage state, registry, and cache. `--keep-user-data` preserves user records and caches. Pip removes the package from the active Python environment. Research projects, writable examples, task environments, run artifacts, agent applications, datasets, and task dependencies are never removed. ## Source Development Contributors use the repository environment rather than the operator install: ```bash uv sync --group dev --extra docs uv run praxist --help ``` See [Contributing](https://praxist.sapient.inc/en/docs/guides/contributing) for verification and [Platform Support](https://praxist.sapient.inc/en/docs/operations/platform-support) for host boundaries. --- # Quickstart Praxist separates installation and configuration from research takeover. The local-terminal lane keeps every setup prompt in the terminal. The agent-managed lane lets Codex or Claude Code perform the same setup decisions in one conversation. Both lanes write the same Praxist configuration and run the same readiness checks; neither selects a project or launches research during installation. They do not share an interaction controller or a second OOBE state file. Check [Installation](https://praxist.sapient.inc/en/docs/getting-started/installation) for host requirements and [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task) for project readiness before takeover. ## Local-Terminal Setup Run the Codex or Claude Code one-line command from [Installation](https://praxist.sapient.inc/en/docs/getting-started/installation#install-and-configure). It installs Praxist and opens the local first-use wizard. To reopen only the wizard later, run `praxist setup --interactive`. An interactive terminal opens the first-use wizard automatically. Use Up/Down and Enter to choose an item. Esc goes back or cancels without inventing a choice. Setup covers these stages: | Stage | Local interaction | Result | |---|---|---| | Install | Pip installs Praxist and its maintained runtime integrations into the selected Python environment. | `praxist` is available; no task-owned dependency is installed implicitly. | | Legal terms | Review the Fair Source License, User Agreement, and data notice in a temporary scroll view, then explicitly agree or cancel. | Acceptance records the exact legal bundle version and digest; it does not enable optional data collection. | | Privacy | When collection is available, separately choose whether to share pseudonymized product-usage status. Nothing is preselected. | Existing consent is preserved; cancellation leaves consent unset. | | Runtime | Choose one setup profile combining an API provider, agent runtime, concrete model, and authentication mode. | A coherent profile is written to the user configuration. | | Readiness | Register skills for the selected agent host, materialize writable examples, and run host diagnostics. | Installation finishes without selecting a project or starting a run. | API keys are entered only in the local terminal. Praxist displays one `*` for each character, supports paste and backspace, and never places the raw value in the command line, shell history, agent conversation, or task project. Esc cancels key entry. ## Agent-Managed Setup Start Codex or Claude Code: ```bash codex --yolo # or claude --dangerously-skip-permissions ``` Then ask: ```text Install and configure Praxist. Follow the packaged OOBE runbook and stop after readiness checks. Do not select a project or start research. ``` The agent reads `docs/agents/oobe-install.md` and uses its structured choices for legal acceptance, privacy, and runtime decisions. It links to the scrollable legal documents and must not infer acceptance or accept on the operator's behalf. An API provider key is never requested in chat. When an API-backed setup profile is selected, the agent pauses for the operator to enter the key through Praxist's local masked prompt. Pip installation is only the package boundary; it does not imply that the operator accepted legal terms or selected a runtime. The agent must continue the OOBE runbook rather than report that setup is complete. Its first state query is: ```bash praxist setup --agent-managed ``` Its JSON identifies the next required decision. The agent reruns it after each decision and cannot claim setup decisions are complete until `setup_decisions_complete` is `true`. Existing credentials, API provider defaults, and doctor readiness never count as an operator profile selection. This lane and the local-terminal lane share only configuration, validation, and recovery contracts. The agent does not emulate terminal keystrokes, and the local wizard does not reproduce the agent workflow. ## Start Research After Reading the Manual Before handing over a project, read [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task). It defines what the project must already provide, what Praxist will add, and which choices the takeover prompt must communicate. Installation success alone is not launch authorization. When the project is ready, start takeover as a separate command: ```bash # Codex praxist --takeover --task-path /absolute/path/to/research-project # Claude Code praxist --takeover --operator claude --task-path /absolute/path/to/research-project ``` ## Choose a Setup Profile The wizard presents coherent setup profiles that bundle an API provider, agent runtime, concrete model, and authentication method. ### Codex-Native Mode: No API Key Choose Codex-native mode to use the existing saved Codex login. This is the shortest first-use path and does not write an API provider key. Only this explicit profile verifies ChatGPT subscription authentication. If the SDK-pinned Codex currently uses an API-key login, setup opens its local interactive login flow and verifies the ChatGPT login before readiness checks continue. Other profiles neither require nor inspect this login. ### API-Backed Profiles: Long Runs For sustained, cost-sensitive research, prefer an [open-source model API](https://praxist.sapient.inc/en/docs/guides/open-source-model-apis) whose cache reuse, quality, and throughput have been checked on a representative workload. The selector includes maintained DeepSeek API, OpenRouter API, and Anthropic API profiles. Enter the selected provider's key at the local masked prompt when requested. Other supported API-backed setup profiles remain available in the same selector. To inspect the exact current profile contract without changing configuration: ```bash praxist setup --list-profiles ``` ## Reopen a Step The OOBE does not create a separate completion marker. It derives current state from the legal-terms acceptance record, existing Praxist configuration, the profile ID recorded by an explicit setup selection, separate product-usage consent, diagnostics, and task artifacts. A repeated setup keeps the current recognized profile selected and never launches takeover. An existing key for the selected API-backed setup profile is preserved without a second prompt. Reopen only the step you need: ```bash praxist setup --interactive praxist --takeover praxist --takeover --operator claude ``` Review or verify the License and User Agreement independently with: ```bash praxist user-agreement review praxist user-agreement status --json ``` Use an explicit project path when the project is not the current directory: ```bash praxist --takeover --task-path /absolute/path/to/research-project ``` Explicit `setup --profile` and provider automation flags remain available for controlled non-interactive provisioning. They do not infer legal acceptance or an operator profile choice. ## Bundled Complete Examples Installation creates writable Python/JAX and Rust reference projects under `~/PraxistExamples` and prints their absolute paths. Praxist never runs against or writes into the read-only copies in its source or package directory. Existing working copies are preserved during upgrades. [Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates) owns the available projects, copy commands, and guidance for choosing a starting point. ## Check Progress Ask Codex or Claude Code in natural language: ```text Report current research progress and list the strongest variant in every completed generation with its task-defined performance metrics. ``` For direct shell operation: ```bash praxist status --json praxist --monitor --latest ``` `Ctrl-C` closes only the foreground monitor. It does not stop the research run. Ask the agent to stop the current run, or use `praxist stop ` directly. ## Next - [Installation](https://praxist.sapient.inc/en/docs/getting-started/installation) defines package and platform boundaries. - [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task) explains project prerequisites and takeover. - [Product Usage Controls](https://praxist.sapient.inc/en/docs/operations/product-usage) explains consent commands. - [`LICENSE.md`](https://github.com/sapientinc/praxist/blob/main/LICENSE.md) is the canonical software license. - [Praxist User Agreement](https://praxist.sapient.inc/en/docs/legal/user-agreement) defines the service terms accepted with it. - [Direct CLI Operations](https://praxist.sapient.inc/en/docs/guides/operators) is the shell lifecycle guide. --- # Your First Task Praxist adapts an **existing runnable research project** into a task project. The original project supplies the executable baseline; the task project tells the generic framework what may change, how evidence is measured, and what counts as credible progress.
```mermaid flowchart LR PROJECT(["Runnable project
code + environment"]) TASK[["Task project
scientific contract"]] PRAXIST["Praxist
generic research process"] EVIDENCE[("Task-local run
evidence + reports")] PROJECT --> TASK --> PRAXIST --> EVIDENCE class PROJECT source class TASK task class PRAXIST system class EVIDENCE artifact ```
## What Must Exist Before Takeover | Required input | Ready means | |---|---| | Research code | The baseline implementation and normal entrypoint are present. | | Runtime | An existing interpreter, environment, container, or remote path can import the required dependencies. | | Data or simulator | Every required asset is locally reachable through the project's normal interface. | | Baseline path | Training, optimization, simulation, inference, or evaluation runs without Praxist. | | Measurable objective | At least one metric distinguishes candidates and its direction is known. | | Scientific constraints | Forbidden changes, validity conditions, and important tradeoffs can be stated. | Prior results and technical documents improve initialization but are not always required. Takeover reports missing prerequisites instead of downloading an unknown dataset, inventing a simulator, or fabricating baseline performance. ## Write the Research Brief The brief is the operator's main scientific input. It should settle: 1. **Objective:** what should improve and which tradeoffs matter? 2. **Evidence:** which metrics, protocol, and maturity level make a result credible? 3. **Execution:** which environment and local assets are allowed, and what is the realistic compute budget? 4. **Exploration:** should literature lookup, the [Deep Innovation Gate (DIG)](https://praxist.sapient.inc/en/docs/guides/deep-innovation-gate), [Quality-Diversity (QD)](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator), and constructive-peer guidance be active? 5. **Operation:** how many peers and generations are appropriate, and may a validated run launch unattended? One or two bounded calibration runs often reveal where a long-run brief needs revision. Readiness and scientific-integrity gates still apply; a brief cannot authorize Praxist to bypass an unresolved contract. ??? example "Detailed takeover brief" Adapt the intent to the project rather than copying the numbers. ```text Invoke `praxist-takeover` in Codex or Claude Code. Use the Praxist checkout at "/path/to/Praxist" on branch "main". Treat the current directory as the existing research project. Verify that its current environment can run the unchanged accelerator-backed training and evaluation path. If no compatible environment exists but every required dependency is locally available, create an isolated task environment without changing the system Python. Create a separate Praxist task project. Configure 12 peers, 30 generations, and a generation duration justified by measured baseline runtime. Disable public literature lookup. Enable QD, enable DIG only for absolute generation zero, and use the constructive-peer ratio as a soft target. If baseline performance is missing and the project already contains everything required to measure it, run a bounded baseline benchmark and record metrics with their provenance. Use the Praxist agent runtime, API provider, and model selected during setup. Define every metric direction, evidence maturity requirement, and protocol-integrity check. Configure durable parent lanes to retain credible Pareto-optimal solutions across genuinely different metric dimensions. Keep partial, diagnostic, suspect, and protocol-failed evidence visible as follow-up signals without treating it as clean parent evidence. Do not download a new dataset or replace the project's existing simulator or runtime. After mandatory evaluator, lane-routing, readiness, and runtime gates pass, launch in detached mode without optional follow-up questions. Report the task path, evidence contract, lane rules, generation-close policy, run ID, and monitor command. ``` ## Start Takeover From the shell: ```bash praxist --takeover --task-path /absolute/path/to/research-project ``` Codex is the default operator interface; add `--operator claude` for Claude Code. In an existing agent conversation, invoke `$praxist-takeover` in Codex or `/praxist-takeover` in Claude Code. Use `praxist-takeover-codex` only for the explicit no-key Codex-native profile. Use `praxist-task-initialization` to create or repair the harness without launching, or `praxist-interactive-task-init` for confirmation-first design. [Agent Skills](https://praxist.sapient.inc/en/docs/user-guide/skills) owns the complete goal-to-skill map. ## What Praxist Adds Initialization adds the smallest practical harness around existing assets: | Harness area | Purpose | |---|---| | Task contract | Objective, scope, permitted changes, metrics, evidence policy, roles, and run settings | | Evaluator and baseline record | One reproducible path from a candidate to structured metrics with provenance | | Retention and close policy | Reachable durable, Pareto, diagnostic, parent, maturity, and generation-boundary decisions | | Resource observation | Unchanged-baseline timing and bottleneck evidence used to plan concurrency | | Task tests | Evaluator output, lane reachability, maturity, resource handoff, and launch readiness | | `experiments/` | Run artifacts outside Praxist source and stable project code | The exact schema and precedence rules live only in [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects). ## Validate Without Starting Takeover performs validation automatically. For direct inspection: ```bash praxist resolve /absolute/path/to/task praxist doctor --task-path /absolute/path/to/task ``` When the task requires ratios, validate a real evaluator summary through the same serializer used in production: ```bash praxist resolve /absolute/path/to/task \ --result-summary /absolute/path/to/evaluation_summary.json ``` ## What Happens During Takeover
```mermaid flowchart LR DISCOVER(["Discover
project / runtime / baseline"]) DESIGN(["Design
metrics / evidence / roles / resources"]) VERIFY(["Verify
task tests / resolve / doctor"]) LAUNCH(["Launch
detached run / status / monitor"]) DISCOVER --> DESIGN --> VERIFY --> LAUNCH class DISCOVER,DESIGN,VERIFY,LAUNCH phase ```
The operator agent: 1. identifies the active Praxist installation, project, execution environment, local assets, technical context, and prior evidence; 2. reuses measured baseline evidence or offers a bounded measurement when every prerequisite is available; 3. turns the brief into task-owned metrics, ranking, protocol integrity, maturity, retention, close, role, prompt, and exploration contracts; 4. observes the unchanged baseline execution path to estimate runtime and the actual resource bottleneck without changing its backend; 5. creates the harness by referencing existing assets instead of cloning the project into Praxist; 6. runs task, evaluator, lane-routing, resource, resolve, and runtime checks; 7. launches only after mandatory gates pass, then reports the task path, run ID, lifecycle state, and monitor command. If a prerequisite or scientific decision is genuinely unresolved, takeover stops at that point and names the missing input. It does not weaken evaluation to make launch succeed. --- # Agent Skills Praxist skills are operator workflows for Codex and Claude Code. Invoke a skill as `$name` in Codex or `/name` in Claude Code. Skill files instruct the agent; they do not run as background services and they do not replace the `praxist` CLI. ## Choose by Goal | Goal | Codex | Claude Code | |---|---|---| | Learn Praxist and inspect host readiness | `$praxist-onboarding` | `/praxist-onboarding` | | Install or repair Praxist runtime dependencies | `$praxist-runtime-install` | `/praxist-runtime-install` | | Build or repair a task harness without launching | `$praxist-task-initialization` | `/praxist-task-initialization` | | Confirm task design interactively | `$praxist-interactive-task-init` | `/praxist-interactive-task-init` | | Initialize and launch with a configured API provider | `$praxist-takeover` | `/praxist-takeover` | | Initialize and launch with a saved Codex login | `$praxist-takeover-codex` | `/praxist-takeover-codex` | | Start, stop, resume, monitor, or inspect runs | `$praxist-control` | `/praxist-control` | | Diagnose run health or generate reports | `$praxist-diagnostic` | `/praxist-diagnostic` | | Gather literature and benchmark context | `$praxist-scientific-research` | `/praxist-scientific-research` | | Draw a terminal line chart | `$terminal-line-plot` | `/terminal-line-plot` | The complete catalog and activation descriptions are generated from the actual skill metadata in [Skills Reference](https://praxist.sapient.inc/en/docs/reference/skills). After installation, read [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task). Then invoke the matching takeover explicitly without having to remember a skill name: ```bash praxist --takeover --task-path /absolute/path/to/research-project ``` For agent-managed installation, the agent follows [the OOBE runbook](https://praxist.sapient.inc/en/docs/agents/oobe-install) and stops after readiness checks. Takeover remains a separate post-manual action. ## Install or Refresh An operator installation registers the bundled skills automatically. Refresh them after upgrading Praxist: ```bash praxist install-skills --target codex --replace praxist install-skills --target claude --replace ``` Run only the command for the host you use. Codex installs under `${CODEX_SKILLS_DIR:-~/.agents/skills}`; Claude Code installs under `${CLAUDE_SKILLS_DIR:-~/.claude/skills}`. This replaces only same-name Praxist skills managed by the current package. Unrelated user skills are not touched. An operator-owned same-name path is reported as a conflict and remains unchanged. Interactive setup can preserve a backup before an explicitly approved replacement. Source contributors can use symlinks for fast iteration: ```bash bash scripts/install_codex_skills.sh bash scripts/install_codex_skills.sh --target claude ``` ## Workflow Boundaries - Onboarding inspects and explains; it does not launch research. - Task initialization writes only the selected task project. - Control owns lifecycle operations and uses the CLI. - Diagnostic is analysis-only unless the user explicitly asks for task-level improvement. - Scientific research records sourced context; it never claims literature as measured task performance. Task construction and takeover are explained in [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task). Report semantics live in [Run Reports](https://praxist.sapient.inc/en/docs/guides/user-facing-reports-and-init), and lifecycle commands live in [Direct CLI Operations](https://praxist.sapient.inc/en/docs/guides/operators). --- # Task Projects Task projects are research problems that Praxist runs. They are explicit inputs, not bundled system plugins.
```mermaid flowchart LR PROJECT(["Research project
code / environment / assets"]) TASK[["Task project
objective / evaluator / task.yaml"]] PRAXIST["Praxist
generic orchestration"] ARTIFACTS[("experiments/
task-local run artifacts")] PROJECT --> TASK --> PRAXIST --> ARTIFACTS class PROJECT source class TASK task class PRAXIST system class ARTIFACTS artifact ```
This page is the sole detailed owner of the task-project contract. Tutorials summarize the workflow and link here rather than defining a second schema. Planning uses Principal Investigator (PI) agents; multi-PI topologies add a Chair that consolidates their proposals. [Panel Topology Prompts](https://praxist.sapient.inc/en/docs/concepts/panel_topology_prompts) owns those roles. ## Location For local dogfood, put task projects under the ignored repository-root `tasks/` directory: ```bash tasks/my_research_task/ ``` For real collaboration, keep the task in its own Git repository and pass its path: ```bash cd /path/to/task-project praxist start --daemonize --json # or pass --task-path explicitly from another directory ``` ## Required Shape A task project normally contains: - `task.yaml` for task identity, workflow selection, metric, plugin refs, and execution defaults; - `description.md` for stable task context; - `roles/` for task-local role skills. Praxist resolves the declared peer RoleSkills before the research loop, injects each peer's agenda-assigned Markdown contract into that peer's prompt, and records the effective role reference plus content hash on the runtime request; - `audit_rules/` for task-local proposal, result, or agenda criteria; these should be declarative YAML or Markdown by default, not Python framework hooks; - `evaluations/` for task-local evaluation logic. Expensive executable tasks should expose one public evaluator command, normally `evaluations//run.py`, so agents do not call internal harness scripts directly; - `assets/` for harness code, optional reference implementations, fixtures, data metadata, and literature packs; - task-local tests for the harness and any optional reference implementations. Hardware planning observes the unchanged baseline and records its actual backend. A task declares only profiles that its public evaluator can use; it does not infer accelerator requirements from host inventory. Process handoff, NVIDIA/CUDA UUID rules, natural-unit parallelism, and supply feedback are owned by [Central Experiment Scheduler](https://praxist.sapient.inc/en/docs/guides/central-resource-scheduler). Use `templates/tasks/template` as replaceable scaffolding. Complete examples serve a different purpose and must run from their installed writable copies. [Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates) owns that distinction and the available starting points. ## `task.yaml` Contract The task descriptor is the only machine-readable source Praxist needs from a task project at startup. A typical descriptor includes: - task identity: stable id, name, version, and description path; - workflow defaults: enabled stages, generation count, cohort size, and run defaults; - metric contract: primary metric name, direction, optional secondary metrics, mature-evidence ratios, constructive target, and cooperative launch guard; - plugin refs: generic workflow stage, agent runtime, API provider, tool, budget, panel, and graph refs; - task-local refs: roles, audit rules, evaluations, and budget profiles; - assets: harness paths, baseline files, literature packs, dataset metadata, and optional reference implementations; - task entrypoints: the public evaluation command and any optional internal harness runners; - output policy: default experiments directory and artifact retention notes. - optional runtime environment: task execution cwd, venv/python path, PATH additions, and non-secret task env vars. - agent reasoning policy: `agent.reasoning_effort` applies one `auto`, `off`, `low`, `high`, or `max` choice across API providers to peers and planning calls; omitted values default to `max`. Runtime-specific mappings are defined once in [Agent Runtimes](https://praxist.sapient.inc/en/docs/guides/agent-runtimes#reasoning-policy). - generation-scoped Deep Innovation Gate (DIG) and Quality-Diversity (QD) policy: independent enable switches, generation scope, candidate-pool requirements, diversity-cell fields, allocation caps, and task-owned allowed/disallowed file rules. - Gems policy: `gems.enabled`, optional reset cadence, compact Gem caps when reset is enabled, relevant lanes, metric keys, and task-owned maturity thresholds for staged or full-coverage evaluation protocols. - Frontier lanes: optional `evaluation.frontier_lanes` entries that keep mature candidates, promising validation candidates, and diagnostics in separate task-owned evidence streams. Without frontier lanes, Praxist keeps the legacy primary-metric frontier but does not materialize an incubator-style lane for preliminary, aligned, or partial signals. Every lane should declare `parent_eligible`: true only for mature durable parent lanes, false for lower-stage and diagnostic lanes. - Research-loop tools: declare the selected `tool_server:*` refs so resolve and runtime agree. [Tool Servers](https://praxist.sapient.inc/en/docs/guides/tool-servers) owns the catalog, [Scientific Literature Lookup](https://praxist.sapient.inc/en/docs/guides/scientific-literature-lookup) owns source policy, and [Run Reports](https://praxist.sapient.inc/en/docs/guides/user-facing-reports-and-init) owns report behavior. - Metric directions: every metric used for frontier ordering, baseline-beat detection, dimension winners, or charts must resolve to an explicit task-owned `maximize` or `minimize` declaration. Result aliases inherit the direction of their configured source metric. Unknown direction remains unknown; reports do not guess that it should be maximized. - QD plan versus evidence: `planned_dimensions` describes allocation intent; `design_dimensions` describes the implementation that actually ran. Praxist never fills missing evidence from the plan. The complete behavior is defined in [Quality-Diversity Allocation](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator). The descriptor may use the `praxist_plugins` block to bind generic plugin refs and task-local component refs. Do not use task descriptors to smuggle system code into core. ## Baseline Records Serious task projects should keep compact baseline evidence under `assets/baselines/`. Prefer: - `results.jsonl` for machine-readable metric rows; - `curated_baseline_summary.md` for human-readable interpretation and provenance; - `baseline_performance_status.md` for measurement status, command, data source, environment, and any missing requirements. If task initialization can measure the baseline on the current machine but no baseline record exists, Codex should ask the operator whether to write explicit zero placeholders or run a task-local baseline benchmark first. A benchmark run belongs under `experiments/baseline_bench_/`, should use the public task evaluator or documented baseline command, and may use bounded parallelism only when the task and hardware make that safe. Zero placeholders must be marked as placeholders, not measured performance facts. Files under `assets/baselines/` preserve evidence and provenance; they do not implicitly configure runtime comparisons. Verified values must also be declared under `task.yaml:baselines` with their metric direction. `praxist resolve` and `praxist start` warn when a conventional parseable result asset is present but that declaration is empty. The warning is advisory and never auto-imports or trusts task files. ## Evaluator Launch Readiness Task initialization proves an evaluator before expensive fan-out. It first exercises the task-appropriate build/load/startup boundary in the actual runtime, validates the evaluator's public CLI, function, RPC, simulator, container, notebook, or service contract as applicable, and then runs a **one-unit canary** through the public evaluator, the central scheduler when Praxist owns launch, and the canonical summary writer. "One unit" is deliberately task-defined: it is the smallest valid case for that task, not a framework-wide seed, epoch, split, iteration, or hardware rule. The resulting summary must validate and project into a finding before wider execution begins. Any implementation or command change requires a new canary. The canary distinguishes broken execution from a valid weak or negative scientific result; it does not establish performance or mature evidence. A task that explicitly claims independently trusted evaluation must additionally provide a task-owned verifier and demonstrate that peers cannot replace the authoritative result and that altered or unattested evidence is rejected. Peer-authored evaluators retain the normal path and do not inherit an external attestation requirement. ## Override Rules Operators can override run settings from CLI or config. Overrides should change runtime choices, model profiles, budget envelopes, generation count, or local paths. They should not mutate the task source. The intended priority is: ```text CLI args > explicit env vars > override spec > task.yaml defaults ``` Credentials are the exception: raw secrets are resolved by Python credential resolution and are never copied into `task.yaml`. ## Experiments Directory Task run outputs belong under a task-local ignored directory, normally: ```text /experiments/ ``` Do not write long-run task outputs into `templates/`, `examples/`, or `praxist/`. Selection order is explicit `--run-dir`, then `$RUN_DIR`, then a timestamped run under `/experiments/`. Praxist rejects destinations inside its own source checkout. Pass `--run-dir` explicitly for an external output root. The detached launcher log is `/logs/launcher.nohup.log`. Task harnesses publish structured result summaries and findings. They do not write frontier, Gems, prompt, report, or memory state directly. Every ranked metric declares its own direction, and negative evidence uses structured valence/failure metadata rather than prose alone. The canonical, validation, derived, audit, and partial artifact roles are defined once in [Architecture](https://praxist.sapient.inc/en/docs/concepts/architecture#state-and-replay). ## Runtime Environment Tasks that need a specific virtual environment or executable can declare it in `task.yaml`: ```yaml runtime_environment: cwd: task_project # task_project, run_dir, or a task-relative path venv: .venv # task-relative or absolute path # python: .venv/bin/python # optional override; inferred from venv otherwise path_prepend: - bin env: TASK_MODE: dogfood # non-secret task vars only ``` At startup Praxist validates the configured paths by default, injects `PRAXIST_TASK_VENV`, `VIRTUAL_ENV`, `PRAXIST_TASK_PYTHON`, `PRAXIST_TASK_SHELL_PREFIX`, and prepends the venv/python directories to `PATH` for agent runtime sessions. Task prompts and harness commands should prefer `$PRAXIST_TASK_PYTHON` when present and otherwise fall back to `python`. When a task interpreter is declared, experiment children do not inherit the Praxist runner's `PYTHONPATH` or `PYTHONHOME`. A task that genuinely requires either variable must declare it in `runtime_environment.env`; Praxist then treats that value as task-owned. This boundary prevents an older task Python from importing packages out of the runner environment while retaining explicit task-specific import layouts. Do not put raw API keys in `runtime_environment.env`; model and tool secrets belong to credential resolution. Keep dependency installation as a separate task setup step. Task-specific non-secret variables, interpreter selection, and import paths belong in `runtime_environment`; start and resume remain owned by the Praxist CLI. ## Protocol Intent Belongs To The Task Owner Praxist does not impose a universal full-protocol-only policy. Task initialization should resolve one protocol-intent table from the current user instruction first, then compatible project evidence, then a proposed default. For every evaluator mode, the table states whether that mode may launch, rank, count as mature, supply a durable parent, and satisfy close. An explicitly requested partial, scout, reduced-coverage, or otherwise incomplete protocol is valid. Its summaries must still report the actual stage, effort, coverage, and integrity facts, and the task's maturity, lanes, Gems, and close settings must all encode the same choice. Praxist should preserve useful signals from other modes without presenting undeclared deviations as mature evidence. Launch validation must inspect structured evaluator modes and output metadata, never reject commands because their text happens to contain words such as `smoke`, `scout`, or `partial`. The configurations below are recommended defaults when the user and project do not specify different semantics. They are not global restrictions. ## Deep Innovation Gate, Quality-Diversity, And Gems Defaults For newly initialized real research tasks, the current recommended default is continuous evolution with DIG limited to absolute gen0, independent QD enabled for both the gen0 DIG pool and later PI synthesis, and periodic Gems reset disabled: ```yaml generation_policy: max_generations: 8 # or another value suitable for a real research run cohort_size: 5 per_generation_hours: 5 evaluation: maturity_policy: min_effort_ratio: 0.75 min_coverage_ratio: 0.80 require_ratio_gate: true complete_stage_labels: [complete] preliminary_stage_labels: [preliminary, aligned] constructive_peer_mix_enabled: true constructive_target_ratio: 0.75 launch_guard: enabled: true # Observed p90 runtimes of ordinary heavy work and the evaluator whose # evidence is authorized for normal close. estimated_heavy_eval_minutes: 0 estimated_close_grade_eval_minutes: 0 safety_factor: 1.25 synthesis_trigger: mature_quorum_fraction: 0.25 quality_diversity: enabled: true initial_generation_enabled: true later_generations_enabled: true max_same_diversity_cell_peers: 1 max_same_mechanism_family_fraction: 0.34 max_same_intervention_surface_fraction: 0.50 gems: enabled: false selection_policy: mature_evidence_top_k # Smoke placeholder; real tasks use their complete-protocol unit count. min_mature_eval_units: 1 evidence_stage_min_units: complete: 1 max_resets: 3 max_gems_per_reset: 4 max_gems_total: 4 max_gems_per_family: 2 prompt_max_gems: 4 archive_ordinary_findings: true ``` The excerpt keeps operator-facing evaluation, QD, and Gems controls visible. Task initialization manages DIG's internal planner settings from the requested enablement, generation scope, and runtime budget. For tasks with task-defined mature/complete evidence, use a positive mature quorum so raw progress or diagnostic findings cannot become normal completion. `mature_supply_fraction` only prioritizes evidence production; it is not a close gate. Use `0.0` only when the task intentionally has no separate close-grade evidence contract and the operator explicitly accepts information-density closing. When mature/complete evidence is required for normal close, the finalized task must satisfy: ```text estimated_close_grade_eval_minutes * safety_factor < effective_generation_close_horizon_minutes - drain_margin_minutes ``` Use the close-authorized evaluator's observed p90 runtime and at least a 30-minute drain margin unless measured publication/shutdown latency requires more. `estimated_heavy_eval_minutes` may separately describe a longer optional protocol; older tasks that omit the close-grade field use the heavy estimate as a compatibility fallback. The effective horizon is the earliest enabled generation or synthesis hard bound, including an enabled adaptive ceiling. Praxist rejects a declared required-evidence contract that cannot finish by construction. A user-authorized reduced or late-signal protocol remains valid when its maturity, lane, launch, and close settings consistently describe that intent. Smoke fixtures may keep `max_generations: 1` while still declaring these blocks so startup and config parsing stay visible; they may set later-generation QD and next-generation constructive feedback to `false` because neither can take effect in a one-generation run. Enable periodic Gems reset only after an operator request or a diagnostic pass identifies a performance ceiling and recommends a reset cadence. `reset_interval_generations` is meaningful only when `gems.enabled: true`. When the user-approved maturity contract uses ratios, each canonical evaluator summary must emit `effort_ratio` and `coverage_ratio` in a supported scalar fact container such as the summary root, `metrics`, `extra`, or `current_aggregate`. `effort_ratio` is actual training/search/optimization effort divided by the task-defined mature reference effort. `coverage_ratio` is completed required evaluation units divided by total required units. Praxist uses one maturity extractor for the source summary and its auto-materialized finding, so task code must not rewrite the same facts into a second artifact. A standalone manually authored result finding with no canonical summary reference must carry the ratios itself. Task-specific stage labels remain audit context. Use `require_ratio_gate: true` only when the declared contract uses these ratios. An explicit user choice may instead use task-owned labels/flags or information-density closing; without any declared maturity facts, maturity remains unknown. Before launch, validate an actual file from the evaluator's canonical summary writer with: ```bash praxist resolve /path/to/task --result-summary /path/to/evaluation_summary.json ``` The check uses the runtime extractor and requires finite effort and coverage ratios only when the task enables the ratio gate. Passing stage labels are not required and cannot make missing ratio telemetry computable. Canonical summaries must resolve to one completion decision under the task-owned policy. Status vocabulary alone is not decisive: a fixed-budget task may define reaching its configured cap as mature completion. The summary must distinguish that case from an early stop through its achieved protocol, effort/coverage, and completion fields. Task initialization tests both outcomes through the real summary writer so the task policy, Frontier, Gems, and prompt-facing views interpret the same result consistently. Tasks with staged or full-coverage evaluation should define task-owned Gems maturity through `selection_policy: mature_evidence_top_k` and `min_mature_eval_units`. They may map task-owned stage labels to cumulative evaluation-unit thresholds with `evidence_stage_min_units`; Praxist configuration always calls these counts units, regardless of any task-local evaluator term. Do not copy one task's threshold or stage labels into another task; derive both from the complete protocol. ## Frontier Lanes And Incubator Evidence `frontier/frontier_manifest.json` is the canonical machine-readable state for the run's Frontier lanes and validation candidates. The manifest may contain `lane_frontiers` when the task configures `evaluation.frontier_lanes`. Operator-facing leaderboards are usually computed from the SQLite finding store or from findings JSON fallback; a missing standalone leaderboard file is not by itself evidence corruption. Use frontier lanes when the task has staged validation or expensive full evaluation: ```yaml evaluation: primary_metric: score direction: maximize frontier_lanes: - name: confirmed description: "Fully scored, promotable candidates." k: 3 cumulative_cap: 10 axes: - {name: score, direction: maximize} include_lanes: [confirmed, performance] require_metrics: [score] parent_eligible: true - name: incubator description: "Lower-admission durable long-term library for task-authorized, protocol-passed, non-suspect Pareto/new-high candidates needing follow-up." k: 8 cumulative_cap: 48 admit_new_high: true axes: - {name: score, direction: maximize} # Put additional distinct metrics here when they should define # Pareto/new-high retention and are emitted by every result mode the # user-owned protocol authorizes for this lane. optional_axes: # Optional axes are secondary sort/display signals only; they do not # by themselves define Pareto dominance. - {name: secondary_tiebreak_metric, direction: maximize} - {name: diagnostic_display_metric, direction: minimize} include_lanes: [incubator, performance] require_metrics: [score] # This default excludes reduced modes. Adapt the structured filters when # the user's protocol intent authorizes one of those modes as a parent. require_falsey_metrics: [is_smoke_eval, partial, scout_only, validation_only, validation_only_result, late_after_generation_boundary, suspect_protocol, suspect_leakage] parent_eligible: true allow_non_promotable: true allow_missing_tier: true allow_risk_violating: true - name: task_candidate description: "Promising preliminary, aligned, or partial evidence retained for validation." k: 5 cumulative_cap: 20 axes: - {name: score, direction: maximize} include_lanes: [task_candidate, candidate] require_metrics: [score] parent_eligible: false allow_lower_tier: true allow_non_promotable: true allow_missing_tier: true - name: diagnostic description: "Controls, falsifiers, negative evidence, and process diagnostics." k: 2 cumulative_cap: 10 axes: - {name: score, direction: maximize} include_lanes: [diagnostic, control, process, reference, negative_control] parent_eligible: false allow_lower_tier: true allow_non_promotable: true allow_missing_tier: true ``` Rename these lanes and metric requirements for the domain. The important contract is not the specific names; it is that a lower-admission incubator keeps task-authorized, protocol-passed Pareto/new-high variants available as long-term parents, while modes marked non-parentable but still useful have a declared candidate lane instead of being forced through the same gate as fully promotable results. `allow_lower_tier: true` retains lower-stage signals for revalidation; it does not make that lane a durable parent source. The lane name in an evaluator summary is a **source label**; each configured lane is a durable target selected by Praxist. When confirmed and incubator should both consider ordinary clean parent-authorized results, the evaluator should normally emit one shared task-owned source label such as `performance`, and both targets should list it in `include_lanes`. Do not map every such result directly to `confirmed`: candidates outside confirmed top-k then cannot reach an incubator that only accepts `incubator`/`performance`. At evaluator ingestion, `frontier_lane`, `promotion_lane`, and `lane` are accepted source-label fields in that precedence order; the first non-empty value is used. Committed entries store the selected target in `frontier_lane` and `promoted_for_lane`, and preserve a different submitted source label in `source_frontier_lane`. Validation-candidate records use `submitted_frontier_lane`. Task initialization must run a task-local reachability regression against real summary construction. For each parent-eligible target, at least one protocol-passed fixture from a parent-authorized mode must satisfy its source and metric filters. With both confirmed and incubator present, test more than `confirmed.k` parent-authorized candidates. When the task has a justified distinct incubator axis, include a candidate that is outside primary top-k but non-dominated on that axis; it must remain incubator-eligible. Do not invent a secondary metric for a genuinely single-metric task, while fixtures from modes the task marks non-parentable, plus protocol-failed, validation-only, late, and suspect fixtures, remain non-parentable. A temporarily empty incubator is valid when no new Pareto point exists; an unreachable incubator is a harness defect. Durable capacity is evidence-based rather than alias-based. Multiple finding or variant names that reference the same exact immutable result artifact (`source_result_path` and SHA-256) consume one durable lane slot. A different path or hash is a different artifact, so independent replications remain eligible. Semantic variant identity and lineage remain separate from this capacity rule. When non-code launch settings can alter a treatment, the evaluator summary should additionally own a secret-free top-level `effective_config` object and `effective_config_complete`. Include every task-owned argument, environment override, protocol choice, and config-file value needed to reproduce the treatment **after** the evaluator has applied defaults, aliases, parsing, and type conversion. The resolved treatment is authoritative: omitting a setting and explicitly supplying its resolved default must produce the same object and digest. A genuinely different resolved value must produce a different digest. Do not use an unfiltered process-environment snapshot as the contract. Praxist carries a deterministic digest and the existing summary path into findings, frontier, Gems, PI context, and reports; it does not copy the full object into those derived views. For derived work, a task-owned evaluator or existing launch helper should read the selected parent summary and compare the child using this same resolved schema before expensive execution. It may inherit allowlisted scientific values when the task defines that behavior, but must not replay a complete parent process environment or expose credentials. For an exact replication, publish `replication_of_effective_config_sha256` from the selected parent. Only a completed result whose current complete digest matches that value supports the exact-replication label. A result without these optional fields remains fully compatible and follows the task's existing maturity and promotion rules. ## Evaluation Entrypoint Task-owned evaluation may use any internal harness layout, but Praxist-facing instructions should name a single public command. Prefer the structured `task_entrypoints.evaluation.command` field; Praxist normalizes it into the legacy-compatible `toolchain.eval_entrypoint` only in memory when needed: ```yaml task_entrypoints: evaluation: command: evaluations/pareto_tiered/run.py output_policy: compact stdout plus raw evidence under the run directory ``` Agents should call the evaluation entrypoint. Internal benchmark files under `assets/harness/` are task implementation details and may change without changing the Praxist task contract. Keep task-owned evaluator, trainer, config, and harness paths relative to the task root. The scheduler resolves a statically identifiable task-owned command entrypoint against that root while preserving the task's configured working directory, even when the operator starts or resumes Praxist elsewhere. Absolute paths remain valid for external datasets, simulators, environments, or services, but task initialization must verify each one before launch. A new task is launch-ready only after the declared interpreter runs the public evaluator successfully from both the task root and a run-like working directory and resolves the same task-owned paths in both cases. Compact result summaries may be nested under `results/**/` and use `summary.json`, `evaluation_summary.json`, `eval_summary.json`, `tiered_eval_summary.json`, or `custom_*_tiered_eval_summary.json`; `result_summary.json` is retained for compatibility. Put task-owned lane, maturity, effort/coverage ratio, protocol, parent-use, and diagnostic metadata in structured summary fields so materialization can preserve it in canonical findings. Publish a stable top-level `variant_id` for each candidate (or an explicit child-result ID when one evaluator emits several candidates) and reuse that identity across evaluation stages; directory names remain fallback provenance rather than candidate identity. Praxist normalizes `task_entrypoints.evaluation.command` into the legacy-compatible `toolchain.eval_entrypoint` when the latter is absent, then forwards that single public evaluator to the default Claude SDK runtime. A direct Bash invocation is registered through the existing protected-PID process-group launcher, so generation drain and `active_evals` observe the same job. Keep the explicit `protected_pids launch` form in task prompts whenever task-owned code launches child work. New tasks using central admission declare resource profiles and submit with a stable semantic tag, profile, and work class. Profiles come from a timestamped trace of the unchanged baseline rather than a synthetic CPU-versus-accelerator comparison or a single teardown sample. Supply defaults, pressure rules, process ownership, retries, and backend-specific handoff checks are defined in [Central Experiment Scheduler](https://praxist.sapient.inc/en/docs/guides/central-resource-scheduler). ## Boundary Rules - Praxist core discovers a task only after the CLI passes an explicit resolved task-project path. The operator CLI resolves `--task-path` > `TASK_PATH` > invocation directory. - Task-local refs such as `task_role:*`, `task_audit:*`, and `task_evaluation:*` are resolved inside the task project, not through the global plugin catalog. - The Praxist repo should not gain task-specific roles, audits, evaluations, or harness code in `praxist/plugins/**`. - A task project may ship its own generic plugins (e.g. a task-specific `panel_topology`) under `/.praxist/plugins///`. These are discovered with `source="task_project"` and take priority over a same-named bundled plugin. They are only scanned when `--task-path` selects the task; no implicit scan of arbitrary task directories happens. ## Task Tests Task projects should include their own tests for the harness, fixtures, evaluation profiles, and optional reference implementations. The Praxist repository tests the task-project boundary and templates; it should not become the permanent test suite for a private external task. --- # Examples And Templates Praxist ships two kinds of task-oriented assets. Their purposes are deliberately different. | Asset | Use it when | What it contains | |---|---|---| | Template | You are creating or testing a new task project | Replaceable scaffolding, placeholders, and deterministic smoke fixtures | | Example | You want to inspect a finished integration | A complete task harness, evaluator, task-specific code, redistributable assets, evidence, and tests | Templates under `templates/tasks/` are meant to be copied and adapted. Their defaults illustrate contract shape; they are not universal scientific policy and may intentionally omit datasets or production evaluators. Examples under `examples/` are runnable reference projects. They retain their own domain assumptions, metrics, resource profiles, and evidence boundaries. Those choices demonstrate one project and must not become Praxist-wide defaults or be copied uncritically into another field. Both remain outside the `praxist` system package. Praxist core and generic plugins contain no facts from either asset. Run lightweight checks in place, but copy a template or example outside the Praxist checkout before starting a research run so generated artifacts remain external to product source. ## Materialize A Complete Example First-use setup creates writable copies under `${PRAXIST_EXAMPLES_HOME:-~/PraxistExamples}`. Inspect or recreate them with: ```bash praxist examples list praxist examples install rocket_booster_recovery praxist examples install rocket_booster_recovery_rust ``` Pass `--destination /absolute/path` to install one example elsewhere. Existing destinations are preserved unless the operator explicitly chooses another path. ## Choose A Starting Point - Start with a [task template](https://praxist.sapient.inc/en/docs/reference/task-templates) when authoring a new task contract. - Study [Rocket Booster Recovery](https://praxist.sapient.inc/en/docs/examples/rocket-booster-recovery) for a Python/JAX integration or [Rocket Booster Recovery (Rust)](https://praxist.sapient.inc/en/docs/examples/rocket-booster-recovery-rust) for an offline native Rust integration of the same research problem. - Read [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects) for the canonical task contract shared by both. --- # Rocket Booster Recovery Rocket Booster Recovery is a complete classical-control example with a frozen six-degree-of-freedom plant, deterministic data banks, a first-contact landing evaluator, task-local research roles, and preserved baseline evidence. It shows how a real research project and its Praxist task harness fit together without moving domain logic into the framework. The source checkout contains the example at: ```text examples/rocket_booster_recovery/ ``` A Praxist wheel exposes the same tree as the package resource `praxist/resources/examples/rocket_booster_recovery/`. ## What It Demonstrates - a complete task harness rather than a scaffold; - one public evaluator backed by a frozen simulation boundary; - task-owned metrics, maturity rules, frontier lanes, roles, and audit rules; - measured evidence kept separate from unmeasured platform claims; - small frozen data banks and manifests protected by checksums; - lightweight contract tests that do not launch a research run. The controller design, physical constraints, metrics, and hardware profiles are specific to this example. They are not Praxist defaults. See [Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates) before adapting any part of it. ## Inspect And Test The source and package-resource copies are inspection-only. Materialize a writable project before creating an environment or running any code: ```bash praxist examples install rocket_booster_recovery cd ~/PraxistExamples/rocket_booster_recovery python3.11 -m venv .venv source .venv/bin/activate python -m pip install -r requirements.txt ./scripts/run_tests.sh ``` The project README owns the scientific protocol, environment details, baseline measurements, and evaluator commands. ## Start A Research Run The first `praxist setup` after pip installation creates a writable copy at `~/PraxistExamples/rocket_booster_recovery` and prints the absolute path. Recreate it explicitly or choose another destination with: ```bash praxist examples install rocket_booster_recovery praxist examples install rocket_booster_recovery \ --destination /path/to/rocket_booster_recovery ``` An existing destination is preserved without replacement. Prepare that working copy's environment, then choose the task harness that matches the host: | Harness | Intended host | Evidence status | |---|---|---| | `task_GPU_server` | A compatible accelerator server | Includes the project's measured server baseline | | `task_PC` | A workstation or laptop | Requires a baseline measured on that machine before research | ```bash cd ~/PraxistExamples/rocket_booster_recovery praxist resolve "$PWD/task_GPU_server" --run-dir "$(mktemp -d)" praxist start --task-path "$PWD/task_GPU_server" --daemonize --json ``` The bundled server evidence must not be relabeled as evidence from another machine. The project README owns hardware details and scientific protocol. --- # Rocket Booster Recovery (Rust) This complete example implements the Rocket Booster Recovery research problem as a native Rust project. It is independent of the Python/JAX example, keeps its Cargo dependencies vendored for offline builds, and includes three Praxist task profiles: | Profile | Intended host | |---|---| | `task_GPU_server` | High-concurrency Linux server; the evaluator remains CPU-only | | `task_linux` | Portable x86_64 or aarch64 Linux host | | `task_macos` | Apple Silicon macOS host with platform-specific baseline qualification | The packaged copy is read-only. Install a writable project before building, evaluating, or launching research: ```bash praxist examples install rocket_booster_recovery_rust cd ~/PraxistExamples/rocket_booster_recovery_rust ``` The project requires Rust 1.85 or newer. Validate the selected profile from the writable copy: ```bash cargo test --release --workspace --locked --offline praxist resolve "$PWD/task_linux" --run-dir "$(mktemp -d)" ``` Then launch through an agent skill or the direct CLI using that task directory: ```bash praxist start --task-path "$PWD/task_linux" --daemonize --json ``` The example's [source README](https://github.com/sapientinc/praxist/tree/main/examples/rocket_booster_recovery_rust#readme) is the sole owner of its scientific protocol, baseline evidence, build audit, platform constraints, and task-profile details. --- # Direct CLI Operations This guide is for operators who want to control Praxist directly from a shell, without asking an agent to perform the lifecycle action. For guided operation, describe the desired action in Codex or Claude Code; it will use the `praxist-control` skill when appropriate. The `praxist` CLI is the canonical shell interface. The Python module entrypoint is a low-level compatibility surface.
```mermaid flowchart LR VALIDATE(["Validate
doctor + resolve"]) RUNNING["Running
start --daemonize"] OBSERVE(["Observe
status / monitor"]) BOUNDARY[["Lifecycle boundary
stop / resume / complete"]] VALIDATE --> RUNNING --> OBSERVE OBSERVE -.-> RUNNING RUNNING --> BOUNDARY BOUNDARY -.-> RUNNING class VALIDATE,OBSERVE interface class RUNNING system class BOUNDARY artifact ```
## Command Map | Goal | Direct command | |---|---| | Reopen first-use runtime setup | `praxist setup --interactive` | | Hand a project to guided takeover | `praxist --takeover --task-path ` | | Check host and runtime readiness | `praxist doctor --task-path ` | | Validate a task without starting | `praxist resolve ` | | Start a detached run | `praxist start --task-path --daemonize --json` | | List or inspect runs | `praxist status --json` | | Inspect one run | `praxist status --run-id --json` | | Open the read-only TUI | `praxist --monitor --run-id ` | | Stop one run | `praxist stop --grace 300 --json` | | Resume a clean interrupted run | `praxist resume --daemonize --json` | Use `praxist --help` for the live argument contract. The generated [CLI Reference](https://praxist.sapient.inc/en/docs/reference/cli) is derived from that parser. The first command reopens setup. The second is a separate, post-manual project handoff described in the [Quickstart](https://praxist.sapient.inc/en/docs/getting-started/quickstart). Neither replaces the task validation or lifecycle commands below. ## Select the Task and Configuration An explicit `--task-path` wins over `TASK_PATH`, which wins over the invocation directory. Prefer the explicit form in scripts: ```bash praxist resolve /absolute/path/to/task praxist start --task-path /absolute/path/to/task --daemonize --json ``` Lifecycle commands load `${XDG_CONFIG_HOME:-$HOME/.config}/praxist/env` by default. When using another configuration file, pass it to every gate and lifecycle action: ```bash praxist doctor --task-path /absolute/path/to/task --config-file /path/to/env praxist resolve /absolute/path/to/task --config-file /path/to/env praxist start \ --task-path /absolute/path/to/task \ --config-file /path/to/env \ --daemonize \ --json ``` Use the same configuration for validation and startup. Explicit process environment values have the highest credential precedence. See [Credentials](https://praxist.sapient.inc/en/docs/guides/credentials) for secret handling and [Agent Runtimes](https://praxist.sapient.inc/en/docs/guides/agent-runtimes) for runtime selection. For a source checkout, run the same CLI from the repository environment: ```bash cd /path/to/Praxist uv run praxist start \ --task-path /absolute/path/to/task \ --daemonize \ --json ``` ??? note "Shared state and host identity" Registry actions record a best-effort host identity so a shared state directory cannot make one machine stop another machine's run. A minimal or rebuilt container without a stable machine ID should set one persistent, host-local `PRAXIST_HOST_ID`. Never reuse that value across distinct hosts. ## Validate and Start ### Configured API Provider Run the inexpensive gates before launching: ```bash praxist doctor --task-path /absolute/path/to/task praxist resolve /absolute/path/to/task praxist start \ --task-path /absolute/path/to/task \ --daemonize \ --json ``` Preserve explicit `--runtime`, `--model-provider`, `--model`, `--cohort`, `--generations`, and `--strategy` choices when required. `--daemonize` lets the run survive the launching shell or Codex session. `--json` makes run ID, PID, run directory, log path, and monitor handoff machine-readable. ### Codex-native Mode Use the route-aware doctor and pass the same mode to start: ```bash praxist doctor --codex-native --task-path /absolute/path/to/task praxist resolve \ /absolute/path/to/task \ --codex-native \ --runtime agent_runtime:codex_sdk \ --model-provider model_provider:openai_compatible \ --model gpt-5.6-luna praxist start \ --codex-native \ --task-path /absolute/path/to/task \ --agent-system codex_sdk \ --runtime agent_runtime:codex_sdk \ --model-provider model_provider:openai_compatible \ --model gpt-5.6-luna \ --daemonize \ --json ``` The saved-login route is isolated from configured relay API providers and API-key endpoints. Use `praxist-takeover-codex` for the guided equivalent. ## Observe a Run ### Status ```bash praxist status --json praxist status --run-id --json ``` The targeted form is preferable in automation. It avoids mixing unrelated runs when multiple task projects share a host. ### Foreground Monitor ```bash praxist --monitor --run-id ``` The fullscreen TUI is read-only and independent of the research process. It shows run state, peer health/activity, recent log context, and hardware warnings. Visual redraw is decoupled from bounded artifact and hardware sampling, so display responsiveness does not multiply research-side probes. `Ctrl-C` exits only the monitor; the detached run continues. Use `praxist --monitor --plain` for a non-interactive terminal or append-friendly transcript. For peer rows, the live monitor reads each peer's bounded `recent_result_artifacts` summary instead of recursively reconciling the complete result tree in the long-running display. Use `praxist --monitor --once`, `praxist status`, or the diagnostic workflow when you need a complete artifact reconciliation view. `--interval` controls the display interval. Exact defaults and limits are owned by the generated [CLI Reference](https://praxist.sapient.inc/en/docs/reference/cli). ### Interpret Operational State | View | Authority | |---|---| | Result and finding summaries | Measured task evidence | | `frontier/frontier_manifest.json` | Canonical lane and promotion state | | Committed `gems/gems_state.json` | Canonical Gems state | | `gen_N/generation_boundary.json` | Canonical generation boundary | | Leaderboards, Principal Investigator (PI) evidence packs, rendered prompts, reports | Derived views or audit snapshots | | TUI and scheduler status | Live operational telemetry, not scientific evidence | Count completed generations from the contiguous committed boundary markers. If live status or `run_summary.json` reports a larger value than those markers, report the mismatch and classify the extra generation as pending boundary work; do not treat frontier entries or `generation_results.json` alone as a commit. When central scheduling is enabled, `/resource_scheduler/status.json` distinguishes queued, running, blocked, completed, failed, and rejected work. Read lifecycle `running` separately from `running_activity.by_resource_phase`; a live wrapper is not proof of active accelerator compute, and an observation marked `unknown` is not proof of a stall. Resource telemetry helps explain throughput but cannot promote a result. ## Stop a Run Target one verified run whenever possible: ```bash praxist status --run-id --json praxist stop --grace 300 --json ``` After stop returns, poll status until the selected process disappears. Use `praxist stop --all` only when every registered run owned by the environment is intentionally being stopped. Do not use broad `pkill` patterns. The foreground monitor is independent, so stopping a run does not need to find or kill a monitor process. ## Resume a Run For a run stopped at a clean, recognized boundary: ```bash praxist status --run-id --json praxist resume --daemonize --json ``` Never resume a verified live controller. Praxist preserves the original API provider, agent runtime, model, frontier strategy, and task identity; resume rejects changes to these canonical values. An unchanged task checkout may move to a new absolute path; Praxist validates its persisted manifest and effective descriptor. Task identity comes from the persisted task contract. Interrupted final boundaries require more care. Common cases include: - an unfinished final generation after a complete PI agenda; - a finished cohort whose PI/Chair boundary did not finish; - a committed Gems reset followed by a partial next generation; - a pending or incomplete Gems reset transaction. The operator agent should prepare the run directory before calling the resume command for these irregular cases. It should inspect the Praxist resume plan, back up the run before any manual crop, preserve complete generation evidence and committed Gems state, and prefer Praxist's internally recoverable boundary path. When a partial boundary is not recognized, do not hand the partial state directly to `praxist resume`; use a documented repair path or crop to a named clean boundary only with operator approval. Use `praxist-control` with the request "resume the latest run" for this preparation workflow. The skill understands interrupted PI and Gems boundaries and avoids destructive guessing. ## Agent-Assisted Operation Codex or Claude Code is the recommended interface when an action requires path selection, artifact interpretation, irregular resume preparation, or a concise progress report. Ask in natural language, for example: ```text Report current research progress and list the strongest variant in every completed generation with its task-defined performance metrics. ``` The agent will use `praxist-control` as needed. If invoked without an operation, the control skill asks for `start`, `stop`, `resume`, `status`, `monitor`, or `detect-active-runs` instead of guessing. For launch, the agent must know the exact task project before launching. It should confirm the exact path. It should not infer a task from a broad filesystem scan. During a status request it reports generation progress, incubator or leaderboard performance, CPU/memory/process/accelerator load, generated report paths, and at most two score curves through `terminal-line-plot`. It must not stop, resume, crop, rerender, or edit files during a status request. Canonical state remains authoritative; derived reports remain audit snapshots. ### Guided Diagnostics and Reports Use `praxist-diagnostic` when the question is why a run is unhealthy or weak, not merely what state it is in. The default diagnostic is analysis-only and may write a report under task `docs/`; it must not edit task code or active run artifacts. It can audit diversity HHI (Herfindahl-Hirschman Index), artifact consistency, resource/runtime friction, a performance ceiling, and the strongest variants. For persistent weakness it can build a chronological agent behavior analysis report. Manual A/B/C run reports put strongest results first, strong-variant lineage second, and run health third. These reports are derived views: canonical state remains authoritative and report snapshots remain audit snapshots. ## Output Locations `praxist start --json` returns the selected run and launcher-log paths; `praxist status --json` resolves registered runs without relying on the current directory. [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects#experiments-directory) owns output placement, run contents, and task-runtime path rules. --- # Run Reports Run reports are human-readable, derived views of canonical Praxist evidence. They explain progress without becoming another result, frontier, or baseline owner.
```mermaid flowchart LR EVIDENCE[("Canonical run evidence")] REPORT["Markdown + PDF report"] READER(["Researcher / diagnostic agent"]) EVIDENCE --> REPORT --> READER class EVIDENCE artifact class REPORT system class READER actor ```
## Automatic Reports Praxist writes reports under `/docs/praxist_reports/` when: - the first credible frontier result beats a known baseline on the same metric; - the completed-generation count reaches a multiple of three (3, 6, 9, and so on); or - the run reaches a terminal state. The generated subtree is excluded from task identity, so report refreshes do not change the task manifest or block a compatible resume. ## Report Structure Every report follows the same order: 1. **Strongest variants or Pareto front** presents task-declared metrics, credibility, and concise mechanism summaries. With no clean frontier entry, dimension winners from lower-authority evidence may appear only as clearly labeled signals. 2. **Strong-variant lineage** follows only those variants through parents, generations, source results, and inherited ideas. 3. **Run health** summarizes artifact consistency, operational friction, diagnostic coverage, and caveats. The PDF companion uses the same facts and adds charts only when task-declared numeric directions support them. A metric with unknown direction may be shown as context, but it cannot select a winner, trigger a baseline claim, or drive a directional chart. ## Generate a Report `tool_server:run_report` exposes `generate_run_report` inside an agent session. The `praxist-diagnostic` skill can generate the same report during an explicit run-health analysis. Canonical truth remains in result and finding summaries, `frontier/frontier_manifest.json`, committed `gems/gems_state.json`, and `gen_N/generation_boundary.json`. Reports never feed promotion or generation close decisions. --- # Cost Estimation Praxist can spend tokens, task-selected compute time, wall-clock time, tool quota, and external API quota. Cost estimates are advisory, not a replacement for BudgetPolicy. ## Inputs To Estimate Estimate cost from: - cohort size; - number of generations; - model profile per stage; - expected peer session count; - expected Principal Investigator (PI) and Chair planning calls; - prompt size and cache stability; - tool calls; - evaluation runtime; - platform/backend capacity and queue behavior, including accelerator capacity only when the task actually uses one. ## Prompt Cache Readiness PromptLayout V1 keeps frozen and dynamic prompt blocks separate. Stable frozen prefixes improve the chance that caches managed by the agent runtime or API provider are useful. Agent runtime cache behavior is currently treated as runtime-managed. Praxist records layout hashes and cache provenance rather than injecting raw cache directives where the runtime does not expose them. ## Interpreting Usage Exact token or cache usage depends on what the agent runtime/API provider returns. Missing metering should be recorded as unknown rather than zero. Use run artifacts, budget ledgers, and API provider invoices together when analyzing cost after a dogfood run. For runtimes that report cache usage, `input_tokens` is the inclusive logical input total and `cached_input_tokens` is the cache-read subset. Calculate: ```text uncached_input_tokens = input_tokens - cached_input_tokens - cache_creation_input_tokens cache_hit_ratio = cached_input_tokens / input_tokens sessions_per_peer_generation = peer session count / peer-generation count ``` Treat an unreported cache-creation value as zero. If the reported components do not fit inside inclusive input, keep the raw values and mark them inconsistent instead of forcing the equation to balance. Claude SDK telemetry additionally preserves cache-creation input separately. Its API provider's native `input_tokens` value is uncached input, so the adapter normalizes inclusive input as uncached + cache read + cache creation. Historical or third-party records whose cached input exceeds their declared total are kept unchanged and marked `telemetry_inconsistent`; Praxist does not clamp them or publish a misleading cache-hit ratio. Treat these as separate signals. Prompt caching can reduce billed compute while logical input remains high; reducing unnecessary fresh sessions reduces both repeated tool work and logical input. The read-only diagnostic inventory derives these values from canonical `generation_results.json` rows and the run summary; it does not create another usage ledger. Session-reuse and tool-output mechanisms are defined in [Cost Optimization](https://praxist.sapient.inc/en/docs/guides/cost-optimization); this page only defines how to estimate and interpret cost. --- # Architecture Praxist is a task-agnostic autonomous-research control plane: ```text stable core contracts + generic plugins + explicit external task projects ```
```mermaid flowchart LR TASK[["Task project
domain truth + evaluator"]] ENTRY(["CLI + resolver
frozen run configuration"]) ENGINE["Praxist
core contracts + generic plugins"] RUN[("Task-local run
canonical evidence + derived views")] TASK --> ENTRY --> ENGINE --> RUN class TASK task class ENTRY interface class ENGINE system class RUN artifact ```
## Ownership Boundary | Owner | Responsibilities | Must not own | |---|---|---| | Praxist core | Stable protocols, resolution, canonical storage, replay, credentials, budgets, and extension interfaces | Task facts, provider wire objects, or workflow implementation details | | Generic plugins | Replaceable runtimes, API providers, workflow stages, tools, graph maintenance, topology, and budget policy | Facts or prompts usable only by one task | | Task project | Objective, baseline, evaluator, metrics, evidence policy, roles, prompts, assets, and scientific constraints | Praxist source or another task's state | | Run directory | Frozen configuration, artifacts, evidence, lifecycle state, and replay records for one run | Mutable task configuration as current truth | The only system package is `praxist`. A component is generic only when two unrelated task projects can use it without importing one task's private facts. The complete task contract is defined in [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects); templates and complete examples are classified in [Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates). Run artifacts never belong in the Praxist checkout. [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects#experiments-directory) owns output-path selection and the task-local run boundary. ## Core and Plugin Boundary `praxist.core` supplies protocols and interfaces for task resolution, runtime requests and events, model profiles, tools, artifacts, findings, trajectory, budget, credentials, workflow stages, topology, and replay. Selectable behavior belongs behind one of those interfaces. Executable generic plugins live under `praxist/plugins/**` or an explicitly selected plugin root. Their manifests identify compatibility, entrypoints, and source-hashed code or assets. Task-local generic plugins can live under `/.praxist/plugins/` and are visible only when that task is selected. [Generic Plugins](https://praxist.sapient.inc/en/docs/guides/plugins) owns manifest and testing rules. [Configuration Discipline](https://praxist.sapient.inc/en/docs/concepts/config_discipline) owns configuration ingress and precedence. [Panel Topology Prompts](https://praxist.sapient.inc/en/docs/concepts/panel_topology_prompts) owns the prompt-asset loader contract. ## Runtime and Workflow Boundaries Agent runtimes consume normalized requests and return normalized events/results. API provider plugins describe API shape, models, credentials, and route capabilities; core calls neither a concrete SDK nor provider directly. [Agent Runtimes](https://praxist.sapient.inc/en/docs/guides/agent-runtimes) and [API Providers](https://praxist.sapient.inc/en/docs/guides/model-providers) own those contracts. `workflow_stage:research_loop` owns the executable research loop. Its topology sidecar records each generation's cohort, while the module API exposes read views and records structured external requests. The bundled executor does not apply queued topology mutations. The detailed boundaries live in [Workflow Stages](https://praxist.sapient.inc/en/docs/guides/workflow-stages) and [Research Topology Audit API](https://praxist.sapient.inc/en/docs/guides/research-topology-and-module-api). [Central Experiment Scheduler](https://praxist.sapient.inc/en/docs/guides/central-resource-scheduler) owns experiment admission and resource allocation inside the research-loop plugin. ## State and Replay Run artifacts have one of five roles: - **`canonical_state`** is current machine-trusted state, including measured result/finding evidence, committed frontier and Gems state, and generation boundaries. - **`validation_signal`** is useful non-durable evidence retained for repair, diagnosis, or follow-up. Task policy decides whether later canonical evidence may promote it. - **`derived_view`** is a regenerable bounded view, such as a leaderboard or run report. - **`audit_snapshot`** records what a stage or agent saw, including rendered prompts and evidence packs. - **`partial_output`** is interrupted or rejected output and is ignored by normal runtime readers. The rule is fewer fact owners, not fewer signals. Planning, promotion, reset, resume, and close read canonical state plus explicitly eligible signals. Derived views and audit snapshots explain decisions but never override their sources. `gen_N/generation_boundary.json` is the commit acknowledgement for a generation. Results without a contiguous boundary remain pending work. A final evidence cutoff prevents late files from rewriting a closed generation; they remain visible as follow-up signals. Auto-materialized findings are idempotent, rebuildable projections of result summaries rather than a second result owner. The [Research Loop](https://praxist.sapient.inc/en/docs/guides/research-loop-variant-generation-flow) owns the write order and inheritance sequence. ## Finding Graph The finding graph is advisory research context. Graph maintainers write edges, health, and compact guidance; query tools expose those views to peers and panels. The graph cannot rewrite raw findings, frontier membership, or task ranking. Alternative graph implementations belong behind the graph-maintainer plugin contract. ## Result Preservation Praxist favors survivable long-running research: ```text capture first, label uncertainty, continue when safe ``` Useful output is retained unless continuation would corrupt the fact chain, expose secrets, damage existing results, or exceed an approved resource envelope. Weak provenance is labeled; missing usage is `usage_unknown`, never silently zero. ## Verification Boundary `AGENTS.md` is the contributor contract and the generated reference is the enumerated API surface. Default tests remain offline: they do not require real model keys, network, GPUs, external task repositories, or long research runs. --- # Research Loop Praxist turns one external task project into successive generations of candidate implementations and measured evidence. Core does not compile a task prompt into code: peers inspect the baseline, implement hypotheses, run the task-owned evaluator, and publish evidence. Praxist coordinates that work and commits what later generations may inherit. Principal Investigator (PI) agents propose next-generation work; multi-PI topologies add a Chair that consolidates those proposals. The [Glossary](https://praxist.sapient.inc/en/docs/about/glossary) gives compact definitions, while [Panel Topology Prompts](https://praxist.sapient.inc/en/docs/concepts/panel_topology_prompts) owns the panel contract.
```mermaid flowchart LR PLAN[["Task contract +
committed agenda"]] RESEARCH["Parallel peers +
task-owned experiments"] EVIDENCE[("Results + findings +
retention lanes")] PANEL["PI agents / Chair +
next agenda"] PLAN --> RESEARCH --> EVIDENCE --> PANEL PANEL -.-> PLAN class PLAN task class RESEARCH,PANEL system class EVIDENCE artifact ```
![One Praxist generation: inherited state feeds design allocation and parallel peers, whose artifacts are externally evaluated into findings, then synthesized into a frontier update, next agenda, and compressed memory that generation g+1 inherits.](../assets/figures/praxist-generation-loop.svg)

Local experimentation becomes global evidence and a committed plan for the next generation.

## 1. Resolve and Freeze the Run Startup resolves the selected task, plugins, prompts, baseline references, API provider, agent runtime, and initial durable state. It writes a run-local frozen configuration before cohort execution. The task schema and precedence rules are defined in [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects); the research loop consumes that contract without adding domain meaning. ## 2. Build the Generation Context `GenerationLoop` combines: - the task prompt, peer role, and allowed work surface; - the committed agenda and optional Deep Innovation Gate (DIG) or Quality-Diversity (QD) allocation; - compact frontier, incubator, Gems, graph, and negative-evidence views; - peer-local research memory; and - the task-owned evaluator and protocol contract. Generation zero starts from the baseline and initial task context. DIG may run once before its cohort when enabled; QD is independently selectable. Later generations use the preceding committed PI/Chair agenda. DIG is normally off, while QD can allocate candidate contracts through the existing synthesis path. [Deep Innovation Gate](https://praxist.sapient.inc/en/docs/guides/deep-innovation-gate) and [Quality-Diversity Allocation](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator) own those mechanisms. The workflow records `gen_N/research_topology.json` so worker identities, declared inputs/outputs, and visibility policy remain auditable without changing peer semantics. ## 3. Execute Peer Work Each peer receives a rendered prompt and a normalized runtime request. A typical peer: 1. reads its task, role, agenda, and inherited evidence; 2. states a mechanism hypothesis and intended evidence stage; 3. creates an independent variant under `variants/`; 4. changes only permitted files; 5. submits evaluation through the task's public evaluator path; 6. writes structured output under `results/`; and 7. publishes a finding with caveats and follow-up. The selected resource policy controls experiment admission. It does not choose the hypothesis or change scientific validity. See [Central Experiment Scheduler](https://praxist.sapient.inc/en/docs/guides/central-resource-scheduler). ## 4. Materialize Evidence A **result summary** answers what the evaluator measured. A **finding** records how later research may interpret and use that result. Praxist recursively discovers recognized summaries and idempotently materializes their usable metadata into canonical findings. Findings preserve actual evidence stage, maturity ratios, metrics, mechanism, caveats, lane intent, parent eligibility, and follow-up. A repeated finding ID can refresh changed non-empty fields without erasing useful older fields omitted from the update. Committed lane membership lives in `frontier/frontier_manifest.json`. A leaderboard is only a derived view. Task-owned lane, maturity, and Gems policy is defined in [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects#frontier-lanes-and-incubator-evidence). Immature scout or partial output remains a validation signal, not incubator content. Artifact-role definitions are owned by [Architecture](https://praxist.sapient.inc/en/docs/concepts/architecture#state-and-replay). The loop plans from canonical state, keeps useful validation signals visible, and never treats old prompts, reports, or evidence packs as current truth. ## 5. Close and Commit the Generation After admitted work drains, the boundary performs one ordered commit: 1. ingest peer findings and result summaries; 2. run a final idempotent evidence refresh; 3. update canonical finding, graph, frontier, incubator, and optional Gems state; 4. refresh peer memory and negative-evidence summaries; 5. build the PI evidence view and synthesize the next agenda; and 6. write `gen_N/generation_boundary.json`. The marker is the completion fact. Results or frontier files without a contiguous marker are pending boundary work for resume to finish. A recorded evidence cutoff makes retries deterministic. Results published after that cutoff remain visible as late validation signals, but cannot enter the closed generation through retry timing. Atomic files visible before the cutoff remain eligible even if ingestion observes them during reconciliation. [Research-Loop Flexibility Controls](https://praxist.sapient.inc/en/docs/guides/research-loop-flexibility-controls) owns close eligibility, mature quorum, drain, and bounded-liveness behavior. ## 6. Synthesize and Inherit PI/Chair synthesis reads committed evidence, prior agenda, task constraints, memory, and diversity diagnostics. It writes the next agenda under `agendas/research_agenda_gen.yaml`. An agenda assigns planned work; it is not measured evidence. Evaluator results and committed retention state continue to own scores, maturity, protocol status, and parent eligibility. A failed or uncommitted agenda cannot drive another cohort. A later peer may restart from the baseline, inherit a durable candidate, repair a credible signal, ablate a strong result, combine compatible mechanisms, or investigate an anti-mainline direction. Praxist supplies evidence and constraints, not a fixed code-generation template. ## 7. Audit the Flow | Question | Canonical location | |---|---| | What was implemented? | `variants//` | | What was measured? | `results//` | | What evidence was published? | `findings/`, `shared_findings/`, or the canonical finding store | | What was durably retained? | `frontier/` and `gems/` | | What did the panel plan next? | `agendas/` | | Did the generation commit? | `gen_N/generation_boundary.json` | A variant present only under `variants/` or `results/` may not influence later planning if no usable finding can be published or materialized. --- # Research-Loop Flexibility Controls These controls preserve useful signals while making maturity, retention, and generation close follow the task owner's declared protocol. Exact task fields and the recommended combined profile are defined in [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects#deep-innovation-gate-quality-diversity-and-gems-defaults). ## Mature Evidence When a task enables ratio-based maturity, its canonical evaluator summary emits: - `effort_ratio`: actual evaluation effort divided by the task's mature reference effort; - `coverage_ratio`: completed required evaluation units divided by all required units. Praxist copies these normalized facts into auto-materialized findings. A standalone finding without a canonical summary reference must carry them itself. With `require_ratio_gate: true`, missing or non-finite ratios remain unknown; stage names cannot fill the gap. A task that deliberately uses labels or completion flags instead may leave ratio gating disabled and define those semantics explicitly. Initialization validates one real file from the evaluator's production summary writer with `praxist resolve --result-summary`. During a run, a durable-looking result missing required ratios remains visible as a validation signal and produces one bounded warning; it is not promoted or counted for mature close. Gems, frontier lanes, reports, and close all consume the same maturity decision. Praxist counts generic evaluation units and never gives a domain-specific stage name global meaning. ## Generation Close `synthesis_trigger.mature_quorum_fraction` controls normal close when a task distinguishes close-grade evidence: | State | Behavior | |---|---| | Positive quorum reached | Freeze new work, drain active work, synchronize evidence, and close normally. | | Positive quorum missing at assessment | Fence ordinary admission while deadline-safe mature top-ups remain eligible. | | Quorum `0.0` | Information density may close normally; valid only when this is the task owner's intended protocol. | | Safety bound or fully drained cohort | Preserve liveness and record insufficient maturity where applicable. | The scheduler's mature supply target is advisory and cannot replace this gate. When close begins, `CLOSING_SIGNAL` blocks every new experiment while already running protected work finishes. A bounded drain grace lets agent sessions publish final evidence before `STOP_SIGNAL`; it never kills an active protected evaluator. Runtime-owned stdout files are not completion facts because a successful command may emit no text. Structured runtime completion and exit status own that state. Task-owned progress files remain valid only when their readiness contract is explicit. The committed boundary records close reason, maturity outcome, and optional peer-mix telemetry. `orchestrator_status.json` exposes that compact state to status, monitor, and diagnostics. ## Durable Incubator An incubator lane is a lower-admission, long-term library, not a stricter winner lane. It can retain protocol-authorized, protocol-passed, non-suspect candidates that establish a task-defined Pareto point or new high even when the confirmed lane is full. Confirmed and mature incubator lanes may be parent-eligible. Preliminary, diagnostic, suspect, protocol-failed, and other task-declared non-parent modes remain in validation lanes for follow-up. Optional display/tiebreak axes do not silently become Pareto dimensions. The complete lane schema, source-routing rules, and reachability test are owned by [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects#frontier-lanes-and-incubator-evidence). Promotion rejection summaries remain inside `frontier/frontier_manifest.json`; Praxist does not create a second runtime fact file. ## Constructive Peer Mix When enabled, Praxist estimates constructive solution work versus diagnostic/control work at each committed boundary. The next generation sees the result as advisory feedback, not a quota or execution gate. Disabling the feature stops both calculation and prompt injection, including historical telemetry on resume. This control is independent of the Deep Innovation Gate (DIG) initial innovation-slot policy and Quality-Diversity (QD) allocation. Those switches are described in [Deep Innovation Gate](https://praxist.sapient.inc/en/docs/guides/deep-innovation-gate) and [Quality-Diversity Allocation](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator). ## Launch Freeze The launch guard cooperates with the central scheduler to freeze queued and new submissions before `CLOSING_SIGNAL`. It covers training, evaluation, scripts, shell launchers, and background processes while still allowing result reading, finding publication, and memory updates. Existing protected jobs drain naturally. Tasks may disable the guard only when they explicitly own an equivalent close-safe boundary. Disabling it does not waive timing feasibility or mature close requirements. Heavy-work and close-grade runtime estimates remain separate because a task may permit long optional evaluation while authorizing a shorter protocol for normal close. Central submission, semantic retry, queue ownership, and process-group behavior are defined only in [Central Experiment Scheduler](https://praxist.sapient.inc/en/docs/guides/central-resource-scheduler). --- # Deep Innovation Gate The Deep Innovation Gate (DIG) is a deep-reasoning innovation process that compares mechanism-level alternatives before implementation. DIG is not an experiment loop. It does not train, evaluate, write variants, or encode task metrics. It uses the selected Praxist runtime/API provider and the task's existing prompt, baseline, and file boundaries. ## Generation Scope and Flow The recommended profile runs DIG only before absolute generation zero. Later generations use committed agendas from Principal Investigator (PI) agents and, in multi-PI mode, a Chair; Gems resets do not reactivate DIG. For an enabled generation: 1. build the normal peer context; 2. map the baseline and generate/critique candidate mechanisms with read-only planner tools; 3. validate one selected contract; 4. add that contract as a dynamic prompt block; and 5. launch the ordinary implementation peer. The selected contract identifies the variant, mechanism, intervention surface, rejected alternatives, planned files/changes, expected metric signature, ablation hooks, and fail-fast checks. A material implementation deviation requires an auditable `contract_amendment.yaml`. ## Artifacts ```text gen_/peers//dig/ baseline_mechanism_map.yaml candidate_pool.yaml candidate_reviews.yaml qd_selection.yaml selected_contract.yaml dig_summary.md ``` These are design/audit artifacts, never empirical findings. Result findings may reference their metadata after evaluation. With `generation_scope: initial_only`, their absence after generation zero is expected. ## Retry and Fallback Malformed planner output or an invalid candidate/contract retries within the configured attempt and total-time bounds. Valid phase checkpoints can be reused only when prompt and artifact fingerprints still match. After all attempts fail, Praxist writes `dig_failure_summary.json`. The default then starts the ordinary implementation path, preserving liveness and an audit trail. Strict tasks may disable fallback, accepting that one planning failure can suppress a peer. ## Control Surface Task initialization can enable or disable DIG independently, limit it to the initial generation, bound planning time and candidate breadth, and choose whether planner failure falls back to direct implementation. These controls affect pre-code reasoning only; they do not change task evaluation or evidence. ## Validation The gate requires enough mechanism and intervention diversity, critiques for every candidate, at least one falsifying or diagnostic alternative, and a complete selected contract. Evaluator, data split, and metric-calculation changes are forbidden by default. Lane fit and duplicate checks apply only when their corresponding task policies are active. This validation predicts whether a plan is coherent. It never claims measured performance or makes the contract a parent. ## Relationship to Quality-Diversity and Gems DIG and Quality-Diversity (QD) have independent switches. At generation zero, QD can allocate one validated candidate from each peer's own DIG pool. Disabling that QD path leaves DIG's quality-first local selection intact. Later QD uses existing PI/Chair proposals and does not call DIG or create DIG artifacts. [Quality-Diversity Allocation](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator) owns both allocation paths and their failure behavior. Gems runs after measured evidence reaches a generation boundary. DIG may read Gems as lineage/duplicate context, but cannot create or promote a Gem. New tasks normally start with continuous evolution and enable periodic reset only after an operator decision or diagnosis. --- # Quality-Diversity Allocation Quality-Diversity (QD) seeks a varied set of strong solutions rather than one winner. The term follows Pugh, Soros, and Stanley's [Quality Diversity: A New Frontier for Evolutionary Computation](https://doi.org/10.3389/frobt.2016.00040). Praxist applies that principle to candidate-plan allocation, not evolutionary genotype search. QD is generation-aware and independently configurable. In absolute generation 0 it extends the Deep Innovation Gate (DIG) read-only planning phase. In later generations, where DIG is off by default, it guides the existing Principal Investigator (PI) synthesis path instead of creating another planner or allocator artifact. Multi-PI planning adds a Chair that consolidates proposals. The goal is to keep DIG's rigor while restoring the exploration breadth that a direct no-DIG run can have. ## Problem DIG improves individual peer plans: - each peer maps the baseline before editing code; - each peer generates several mechanism-level candidates; - each candidate receives a structured critique; - each implementation is locked by `selected_contract.yaml`. The weak point is that peer-local selection can converge. If every peer sees the same frontier, Gems, and research agenda, then many peers may select the same obvious family: reward shaping, calibration, risk repair, or the latest Gem lineage. This produces careful contracts, but a narrower generation. The desired behavior is: ```text deep individual reasoning + cohort-level diversity control + no fixed task-specific algorithm quota + no extra experiments during DIG ``` ## Non-Goals Initial-generation QD does not: - change generation semantics; - add an experiment loop; - run training or any task-defined preliminary, aligned, or complete evaluation during DIG; - create a new agent runtime; - encode task-specific metrics in Praxist core; - force every peer into a hand-written algorithm family; - replace Chair or PI judgment. ## Architecture Initial-generation flow: ```text peer context -> DIG candidate generation -> DIG critique -> peer-local QD selection -> selected_contract.yaml -> implementation peer ``` Initial-generation QD flow: ```text all peer contexts -> run each peer's DIG candidate generation and critique concurrently -> collect candidate pools and reviews -> cohort-level QD allocator chooses one candidate per peer -> selected_contract.yaml is updated per peer -> implementation peers launch ``` Each peer still owns its own candidate pool. The allocator never assigns peer A's candidate to peer B. It only decides which candidate from each peer's own DIG pool should become that peer's locked contract. Later-generation Multi-PI flow: ```text completed-generation evidence -> independent PI memos propose experiments and peer contracts -> union of PI proposals is the candidate pool -> Chair applies prompt-guided, soft quality-diversity allocation -> normal research_agenda peer_contracts -> direct implementation peers (no DIG call) ``` Later-generation single-PI flow uses that PI's normal synthesis over findings, frontier, prior agendas, validation signals, and Gems. The PI forms proposals and chooses the final `peer_contracts` in one existing synthesis call under the same soft QD policy. There is no PI-memo union or Chair in this topology. Both paths are the established non-DIG planning path with a compact policy in prompt context. Later-generation QD is intentionally prompt-guided rather than a second deterministic allocator: it creates no planner, contract format, candidate file, or fact artifact. Diagnostics must therefore inspect the final agenda and, for Multi-PI, the existing PI memos; they must not expect a post-gen0 `dig_cohort_allocation.yaml`. When a task declares `evaluation.diversity_dimensions`, QD-enabled PI/Chair planning records the intended value of each applicable axis in `peer_contracts[].planned_dimensions`. This is a plan, not experimental evidence. The peer reports the implemented/evaluated values under the existing finding field `design_dimensions`. Diagnostics derive the planned Herfindahl-Hirschman Index (HHI) from the agenda and realized HHI from findings, then report missingness and drift. Praxist does not copy plans into missing result evidence and does not turn a missing dimension report into a hard execution gate. ## Selection Policy The allocator works over generic descriptors: ```text mechanism_family intervention_surface intent candidate text risk labels peer lane fit local DIG selection known frontier/Gems/sibling signatures ``` It scores candidates with: ```text selection_score = quality_score + lane_fit_bonus + local_selection_bonus + novelty_bonus + target_keyword_bonus - risk_penalty - diagnostic_penalty_when_not_in_diagnostic_slot ``` The gen0 deterministic allocator then applies cohort constraints. Later PI synthesis receives the applicable quality, novelty, lane-fit, risk, target, label-group, and diversity-cap controls as soft allocation guidance: - max peers per exact diversity cell; - max peers per mechanism family; - max peers per intervention surface; - max peers per intent; - max peers sharing an intent, without assuming which task-owned intents are diagnostic; - optional task-defined keyword targets. Keyword targets are generic and task-owned. Praxist core only sees named text groups such as `architecture_or_representation` or `input_feature_use`; a task project chooses the keywords and minimum counts. ## Target Groups Task projects may define soft minimums under the independent policy block: ```yaml quality_diversity: enabled: true initial_generation_enabled: true later_generations_enabled: true target_keyword_groups: - name: architecture_or_representation min_peers: 2 fields: [mechanism_family, intervention_surface, hypothesis, changes] keywords: [architecture, representation, encoder, attention, model_def] ``` `initial_generation_enabled` is an independent disable switch, but gen0 QD still needs an active gen0 DIG scope because its candidate pool comes from DIG. Later-generation QD does not depend on DIG. Targets are not fixed quotas. If no valid candidate matches a target, the allocator records the miss and continues. The generation must stay live. ## Contract Construction If the cohort allocator keeps a peer's local selected candidate, it preserves the existing LLM-authored contract. If the allocator chooses a different candidate from the same peer's candidate pool, it creates a deterministic contract from the validated candidate sketch: - `variant_name` from candidate name and peer id; - `diversity_cell` from candidate signature; - `mechanism_hypothesis` from candidate hypothesis; - `files_to_modify` from candidate implementation sketch; - `allowed_changes` from candidate changes; - standard forbidden changes for evaluator, split, and metric calculation; - implementation steps derived from the sketch; - expected metric signature from the candidate diagnostic prediction; - ablation hooks from the candidate, with a fallback hook if needed. The normal DIG validator still checks the contract. Invalid allocations are not allowed to reach implementation. ## Allocation Failure Behavior QD is conservative: - if initial-generation QD is disabled, DIG uses quality-first eligible selection without duplicate/cell allocation; - if later-generation QD is disabled, single-PI or PI/Chair behavior is identical to the prior non-DIG agenda path; - if a peer has no valid candidate alternatives, it keeps its local DIG contract; - if allocator validation fails for a peer, it keeps that peer's local DIG contract and records the reason; - if all DIG attempts fail for a peer, the existing DIG fallback-to-direct behavior remains unchanged. ## Expected Effects Compared with peer-local DIG, initial-generation QD should: - preserve the stronger mechanism hypotheses and ablation discipline; - reduce repeated same-family contracts in a generation; - allocate at least some peers to architecture, representation, input-feature, off-mainline, or independent exploration when candidate pools support it; - keep diagnostic/control work bounded through the independent gen0 DIG innovation-slot policy; - keep Gems useful without letting every peer inherit the same Gem lineage. Compared with no DIG, the initial-generation QD stage should: - produce more explicit implementation contracts; - reduce first-intuition coding; - make failures easier to interpret; - avoid changing task metrics, evaluator, data split, or baseline contract by accident. ## Test Requirements Tests should cover: - independent DIG scope and initial/later QD config parsing; - default DIG execution only at absolute gen0, including across Gems resets; - cohort allocation preserves one selected candidate per peer; - max same mechanism family is enforced when alternatives exist; - max same intervention surface is enforced when alternatives exist; - target keyword groups are filled when candidates exist; - local DIG contract is preserved when allocation is disabled; - quality-first DIG selection when initial QD is disabled; - later single-PI synthesis and Multi-PI/Chair prompts receive QD policy only when enabled; - later QD does not call DIG or create a separate candidate artifact; - deterministic override contracts pass the existing DIG validator; - generation prompt injection uses the final cohort-selected contract; - fallback-to-direct behavior still works after repeated DIG failure. --- # Peer-Local Structured Memory For Long-Context Continuity This document describes the Praxist mechanism that improves peer continuity across multiple autonomous sessions without carrying raw transcripts forward. ## Design Goal Each peer may span multiple runtime sessions inside one generation. A later session should understand what the same peer already tried, what evidence it created, what sibling peers shared, and where the current hypothesis stands. It should not need a single unbounded chat transcript to do that. The mechanism preserves the useful parts of a long continuous context: - current hypothesis and open questions; - experiment ledger and abandoned branches; - prior session handoff; - relevant new shared findings; - Deep Innovation Gate (DIG) selected-contract state when present; - anti-anchoring prompts that force reconsideration before repeating work. It deliberately does not preserve raw message transcripts in the prompt. ## Session Boundary The research-loop backend updates memory around each peer session: ```text AutonomousAgentLoop._run_session() -> build session_id -> compose base task prompt + peer-local memory block -> execute runtime session -> record structured session result ``` Task-local content remains unchanged; memory supplies execution continuity and audit artifacts only. ## Artifact Layout For each peer: ```text runs//gen_/peers//memory/ peer_state.yaml experiment_ledger.jsonl session_handoff.md seen_shared_findings.json memory_prompt.md ``` `peer_state.yaml` is the compact state card: - current peer identity; - current hypothesis; - open questions; - known dead ends; - active variant; - last session status; - recent result artifacts discovered for the peer. `experiment_ledger.jsonl` is the append-only local ledger: - session id; - concise summary; - success/failure; - duration and tool count where available; - link to the session log; - compact metrics if result artifacts are found. `session_handoff.md` is the human-readable session boundary summary. `seen_shared_findings.json` tracks which shared findings were already surfaced to this peer's session prompt. `memory_prompt.md` records the exact bounded memory block injected into the most recent runtime prompt for auditability. Per-session prompt and manifest snapshots are retained only up to a bounded count; the stable `memory_prompt.md` and `session_prompt_manifest.json` files remain the latest audit pointers. ## Prompt Injection The runtime appends a bounded section titled: ```text Praxist Peer-Local Structured Memory ``` The block contains: - memory discipline requirements; - current peer state; - selected DIG contract snapshot if one exists; - recent peer-local experiment ledger entries; - new shared findings since the last session; - previous handoff summary; - anti-anchoring check. The section is bounded by a character budget. If it grows too large, it is truncated explicitly. This keeps later sessions grounded without allowing memory to become an uncontrolled raw transcript replay. This prompt block is peer-local and session-local. It does not broadcast full sibling peer contracts, does not replay raw transcripts, and does not grow linearly across generations. ## DIG Compatibility When DIG runs, memory includes a bounded view of: ```text runs//gen_/peers//dig/selected_contract.yaml ``` The selected-contract schema and generation scope are defined in [Deep Innovation Gate](https://praxist.sapient.inc/en/docs/guides/deep-innovation-gate). Memory neither changes that contract nor turns it into empirical evidence. ## Shared Findings Refresh The memory layer reads the generation's shared-findings directory and surfaces new JSON findings that the peer has not yet seen. This lets later sessions benefit from sibling peers without depending on full shared transcript replay. The prompt shows only compact metadata: - finding id; - finding type; - producer; - title or summary. After a session ends, surfaced findings are marked as seen. ## Anti-Anchoring Behavior Every injected memory block asks the peer to answer three questions before continuing the same direction: ```text 1. What evidence supports continuing the current mechanism? 2. What evidence suggests pivoting, ablating, or simplifying? 3. What is the cheapest evidence that could falsify continuation? ``` This is intended to preserve the continuity benefits of long context while reducing over-commitment to a stale local idea. ## Task Boundary The mechanism is task-agnostic: - it does not mention domain-specific metrics; - it does not modify any task project; - it does not change benchmark, evaluator, or data semantics; - it does not alter peer count, model routing, Gems, or Frontier ranking. Task-local prompts and role skills remain the right place for domain-specific research instructions. ## Memory Failure Behavior If memory files are missing, malformed, or absent, the runtime initializes a minimal state card and proceeds. A broken memory artifact should not block a peer from executing its assigned research task. Session completion always attempts to write: - a ledger row; - an updated state card; - a handoff note. If the runtime session fails, the handoff and ledger still capture the failure reason where available. ## Expected Effect Compared with raw multi-session replay, this mechanism should: - improve continuity across peer sessions; - reduce repeated dead-end work; - make session boundaries auditable; - improve use of sibling findings; - keep prompt growth bounded; - preserve cross-peer diversity by forcing local anti-anchoring checks. It is a continuity layer, not a new research selector. The generation-level selection logic still belongs to DIG, Frontier, Gems, the Principal Investigator (PI) panel, and its Chair. --- # Panel Topology Prompts A Principal Investigator (PI) agent independently proposes next-generation work. In multi-PI topologies, a Chair compares those proposals and commits one agenda. > **Principle.** A `panel_topology` plugin can ship its own Jinja prompt > templates next to its manifest. The bundled prompts under > `praxist/plugins/workflow_stages/research_loop/backend/multi_pi/prompts/` > are the fallback, not the only choice. This page documents the `topology.prompts_dir` contract for panel topology plugins. It is the override point referenced from `PanelTopologySpec.prompts_dir` in `praxist/core/panel_topology.py`. ## Why The multi-PI backend renders three Jinja templates: | Template | Renderer | Purpose | |---|---|---| | `base.jinja2` | `BasePI.render_prompt` | Round-1 independent PI memo | | `round2_cross_review.jinja2` | `BasePI.run_cross_review` | Round-2 anonymized cross-review | | `chair.jinja2` | `ChairArbiter.render_prompt` | Chair synthesis prompt | The bundled directory is the default loader source. A panel-topology plugin that needs a different collaboration vocabulary can override one or more templates without changing bundled files. `topology.prompts_dir` lets a plugin ship its own prompts directory alongside its topology contract, while keeping the bundled templates as a fallback for everything the plugin does not override. ## Manifest field A panel topology plugin opts in by declaring `topology.prompts_dir` in its `plugin.yaml`: ```yaml schema_version: 1 name: my_panel kind: panel_topology topology: topology_ref: panel_topology:my_panel prompts_dir: prompts/ # relative to this plugin directory modes: { ... } roles: [ ... ] rounds: [ ... ] ``` Layout on disk: ```text my_panel_topology/ ├── plugin.yaml └── prompts/ └── base.jinja2 # overrides bundled; chair.jinja2 etc. fall through ``` ### Path resolution `panel_topology_from_manifest(...)` (in `praxist/core/panel_topology.py`) resolves the manifest value through `_resolve_prompts_dir`: | `prompts_dir` value | Behavior | |---|---| | absent / `null` / `""` | `PanelTopologySpec.prompts_dir = None`; use bundled prompts only. | | relative string (e.g. `prompts/`) | Resolved against the plugin directory. Must point at an existing directory. | | absolute string | Accepted as-is. Must point at an existing directory. | | non-string | `ValueError` at manifest time. | | relative string without a known plugin path | `ValueError` (cannot resolve safely). | | any value pointing at a missing directory | `ValueError` at manifest time. | All `ValueError`s are raised at topology resolution, not at first template lookup. Manifest-authoring bugs surface before any PI starts rendering. ## Loader chain When `PanelTopologySpec.prompts_dir` is supplied, both `BasePI` and `ChairArbiter` build their Jinja `FileSystemLoader` from the search list: ```text [plugin prompts_dir, bundled prompts dir] ``` Jinja walks the list in order, so a plugin can override one template (say, `base.jinja2`) and let the rest fall back to the bundled version. When `prompts_dir` is `None`, the search list contains only the bundled directory. This applies to all three render sites: `BasePI.render_prompt`, `BasePI.run_cross_review`, and `ChairArbiter.render_prompt`. ## What gets threaded where | Layer | What it does | |---|---| | `PanelTopologySpec.prompts_dir` | Frozen `Path \| None` on the resolved topology. | | `panel_topology_from_manifest(..., plugin_path=...)` | Resolves the manifest value relative to the plugin directory; fails fast on missing dirs. | | `legacy_two_round_executor.run_panel` | Resolves the topology once, pulls `topology.prompts_dir`, and threads it to both `instantiate_pi_roles(...)` and `ChairArbiter(...)`. | | `role_bindings.instantiate_pi_roles` | Forwards `prompts_dir` to each PI constructor. | | `BasePI.__init__` / `ChairArbiter.__init__` | Store `self.prompts_dir` and use it when building the loader. | ## Backward compatibility A panel topology that does not declare `prompts_dir` produces `PanelTopologySpec(prompts_dir=None)` and uses the bundled prompts. The bundled `legacy_multi_pi_two_round/plugin.yaml` follows this path. ## Authoring checklist When adding a panel topology plugin that ships its own prompts: 1. Place templates under `/prompts/` (or another directory referenced by `topology.prompts_dir`). 2. Only override the templates you actually need to change. Templates you omit fall back to the bundled versions automatically. 3. Preserve the public template variables consumed by `BasePI` and `ChairArbiter` — overriding the layout is supported, dropping variables silently is not. 4. Add a plugin-local unit test that asserts the override is wired (see `tests/unit/test_panel_topology_prompts_override.py` for a canonical pattern: render with and without the override, check for a plugin-specific token in the output). ## See also - `praxist/core/panel_topology.py` — `PanelTopologySpec`, `panel_topology_from_manifest`, `_resolve_prompts_dir`. - `praxist/plugins/workflow_stages/research_loop/backend/multi_pi/` — `BasePI`, `ChairArbiter`, bundled prompt templates. --- # Central Experiment Scheduler Praxist can make one run-level scheduler the supported launch path for task experiments. Peers submit task-defined work; they do not choose devices or create scheduled task processes themselves. ## Why It Exists Resource arithmetic is simple only when one component owns four facts: 1. which scientific experiment was submitted; 2. which attempt may start; 3. which process group and physical accelerator were assigned; 4. when that allocation is released. The scheduler owns queueing, final `Popen`, process environment, infrastructure retry, and release. Frontier promotion, evidence maturity, [Deep Innovation Gate (DIG)](https://praxist.sapient.inc/en/docs/guides/deep-innovation-gate), [Quality-Diversity (QD)](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator), Principal Investigator (PI) and Chair planning, Gems, and generation policy remain separate. ## Task Contract ```yaml compute_budget: resource_scheduler: mode: central initial_concurrent_experiments: 2 min_concurrent_experiments: 1 max_concurrent_experiments: 8 supply_signal_enabled: true supply_idle_samples: 3 supply_lease_seconds: 600 mature_supply_fraction: 0.25 mature_supply_redundancy: 3.0 mature_assessment_min_completion_probability: 0.25 exploration_reserve: 1 infrastructure_retries: 1 default_profile: gpu_work profiles: cpu_ordinary: accelerator: cpu pressure_domains: [cpu, memory, io] gpu_work: accelerator: gpu gpu_count: 1 gpu_memory_gb: 20 gpu_utilization_pct: 45 pressure_domains: [cpu, memory, io] ``` Task initialization obtains these values by running the unchanged public baseline and observing it externally. It must not create artificial CPU-only and GPU-only rewrites. Observation is a timestamped process-lifetime series, normally sampled every 100-200 ms. Short tasks are safely repeated or replaced by a longer unchanged representative unit; a teardown `0%` sample or fewer than ten useful samples does not establish zero GPU demand. Utilization uses a robust upper estimate of full-lifetime means, while VRAM uses the observed peak plus modest headroom. Ambiguous or undersampled demand remains unknown and therefore exclusive. The default profile matches the public evaluator's normal resource shape because runtime-assisted submission may omit an explicit profile; ordinary analysis commands should not enter the experiment queue. CPU profiles do not reserve cores per experiment; the operating system shares CPU time and Praxist changes total experiment concurrency from live pressure. A profile that declares CPU, memory, or I/O pressure is not newly admitted while that declared domain exceeds its high-pressure threshold; gradual concurrency adjustment remains the recovery path once pressure falls. GPU profiles use two independent per-device hard limits: declared utilization plus observed external load remains at or below 100%, and declared peak memory plus observed external memory remains below 95% of physical VRAM. A settled job's declared envelope is retained across setup, CPU, accelerator, evaluation, and waiting phases because a later phase may return to its peak. Driver activity is reported separately for diagnosis; it does not authorize transient oversubscription. Feasible devices are ordered by the tighter of their compute and memory headroom. Omit either GPU demand field when demand is unknown; the profile then receives exclusive placement. Omitted scheduler fields use documented defaults. In `mode: central`, explicit unknown keys, misspelled booleans, malformed numbers, profiles, or pressure domains fail during task resolution instead of silently changing scheduling policy. Valid numeric values outside a bounded policy range retain the documented normalization behavior. CPU and GPU profiles are not interchangeable unless the task explicitly says they are scientifically equivalent. Praxist never changes a failed GPU experiment into CPU work implicitly. Omitting `--profile` selects `default_profile`; an explicit unknown profile is rejected instead of silently using another device class. ## Mature Evidence And Idle Supply Feedback The scheduler adjusts capacity, but it cannot invent scientific work. When an open generation has unused concurrency slots, the queue cannot fill them, and consecutive host samples show headroom in task-declared pressure domains, it writes short-lived directed leases under `gen_N/resource_supply/`. Completed peer sessions register as idle and watch only their own lease file through the existing event-driven loop. The evidence controller uses the task's existing maturity policy and canonical result store to maintain: ```text Q = max(1, ceil(cohort_size * mature_supply_fraction)) M = unique mature results already published for this generation D = max(0, Q - M) A_target = min(cohort_size, ceil(mature_supply_redundancy * D)) ``` When a task deliberately configures a larger hard mature close quorum, that larger target replaces `Q` for first-wave and debt supply; otherwise the close contract could never receive enough mature work. In that mode `M` is distinct mature peers, matching the existing peer-quorum semantics; without a hard peer quorum, `M` is unique mature results. Setting mature supply fraction or redundancy to zero disables maturity-priority supply even when a hard close quorum exists. The inverse is also important: a positive mature supply fraction does not create a hard close gate. Tasks that distinguish close-grade evidence must set a positive `synthesis_trigger.mature_quorum_fraction`; otherwise raw information density can normal-close the generation while maturity debt remains. Queued/running mature semantic experiments and outstanding mature-priority leases count toward `A_target`; retries retain one semantic identity. The default `0.25` and `3.0` values are calibrated general defaults rather than domain truth. Setting either value to zero disables mature-priority supply without disabling ordinary idle backfill. During assessment, ordinary admission stops while mature top-ups remain eligible when their calibrated probability of finishing before the generation deadline is at least `mature_assessment_min_completion_probability`. The compact log-normal calibration uses successful wall time divided by declared ETA, with a neutral prior that avoids overconfidence from the first few jobs. Before assessment, existing deadline admission remains unchanged. An unknown ETA remains unknown rather than being converted into false precision. Each lease names currently admissible profiles, carries `mature` or `frontier_followup` priority, and expires if unused. `supply_lease_seconds` sets the bounded response window (default 600 seconds, normalized to 180-3600). Expiry limits when the peer may submit an existing plan; it does not limit the runtime of an experiment admitted before expiry. A peer may respond with at most one already justified experiment selected by the current research plan, evidence priorities, and exploration commitments. The signal selects an evidence class but does not choose a hypothesis, create variants, or weaken evaluation standards, expose a device assignment, or bypass the central queue. The scheduler records a short-lived host-wide capacity claim so concurrent runs cannot promise the same slot; the final launch still performs normal admission and assigns the actual device UUID. A real queued job atomically preempts its own or physically conflicting speculative claims, so idle feedback cannot delay submitted research while unrelated CPU-only runs remain independent. Once published, the claim remains stable across later pressure samples until it is consumed, expires, is atomically preempted by real work, the generation closes, or the run stops. Final launch always rechecks live pressure, so this bounded response stability does not bypass admission. `supply_signal_enabled` disables this feedback, while `supply_idle_samples` controls its consecutive-sample requirement. Outstanding leases count against supply capacity, so N free slots wake at most N peers. A peer that declines an unused lease enters a bounded same-priority exponential cooldown rather than losing eligibility for the rest of the generation. A new experiment submission, a changed priority, or a new generation resets that backoff; retained idle registrations allow later maturity debt to wake the peer again without a long polling delay. `resource_supply.stats` reports `conversion_rate=consumed/granted` and attributes known-priority counts under `by_priority.mature` and `by_priority.frontier_followup`. Unused offers terminate as `declined`, `expired`, or `revoked`; a submission carrying an already expired locator is `stale_submission`, while `reuse_ignored` is reserved for a lease that was genuinely consumed once. These are operational facts only and do not replace mature-result quality or quantity. Grant and terminal transitions are durable before they affect the live lease; restart replays the same event ledger and revokes any grant interrupted before a response, so run-wide conversion does not reset with the scheduler process. If host claim release is temporarily unavailable, the lease remains visible with `release_pending: true`, is excluded from actionable maturity commitments, and is retried by reconciliation instead of becoming hidden capacity. At the first wave, up to `Q` peers receive direct-mature advice while at least one peer retains exploration when the cohort has multiple peers. Once mature commitments satisfy `A_target`, spare leases return to Pareto-relevant follow-ups and then already planned scouts. Scientific selection remains with the research loop. This mechanism is resource-type neutral. CPU, memory, and I/O pressure come from live host observations; accelerator placement remains governed by measured memory/utilization profiles. Simulator instances, licenses, remote services, and other bounded resources remain task-owned limits expressed in the evaluator or through a conservative global experiment cap. The goal is to keep the measured bottleneck supplied without forcing every resource to a fixed utilization percentage. ## Natural-Unit Parallelism Task harnesses should identify independent complete-evaluation units such as seeds, folds, scenarios, simulator instances, datasets, benchmark cases, or restart trials. A multi-accelerator profile is valid only when the evaluator actually distributes those units across every assigned physical GPU UUID, preserves binding through all descendants, aggregates independently of completion order, and drains the complete process group. Declaring `gpu_count > 1` without that implementation does not create parallelism. Do not count the same work both as multiple top-level scheduler experiments and as internal child units of one experiment. Prefer the smallest unit that is independently valid, retryable, and aggregatable without changing the scientific protocol. A long wrapper that serially mixes setup, accelerator work, CPU evaluation, and waiting remains one lifecycle job. Where it cannot be split safely, expose monotonic task-owned progress. A task may fail fast after repeated identical infrastructure or implementation errors prove the remaining units non-runnable, but must retain a structured failure summary. Low scores, negative findings, and heterogeneous scientific failures are not fail-fast conditions. ## Submission ```bash PYTHONPATH="$PRAXIST_WORKSPACE_ROOT${PYTHONPATH:+:$PYTHONPATH}" \ "$PRAXIST_RUNNER_PYTHON" -m praxist.plugins.workflow_stages.research_loop.backend.protected_pids launch \ --run-dir="$PRAXIST_RUN_DIR" \ --peer="$PRAXIST_PEER_ID" \ --tag= \ --profile= \ --work-class= \ --eta= -- ``` The tag identifies science, not execution syntax. Variant/protocol/data coverage/seeds/tier belong in it; timestamps, output paths, retry numbers, logging flags, and harmless command spelling do not. Repeated submissions with the same identity share one queued, running, or completed job. Only exit code 75 is an automatically retryable infrastructure failure. A corrected request whose existing job is `failed` or `rejected` must use the same scientific tag with `--retry-terminal`; this creates a new attempt under the existing semantic identity. An identical retransmission of an already accepted retry remains idempotent while queued, running, completed, or `drained_unknown`; a changed request with the flag is rejected in those states. Without the flag, a terminal duplicate is reported explicitly instead of silently rerunning. Do not append arbitrary retry text to a scientific tag. Pre-launch admission timeouts and transient accelerator-probe rejections are the exception: no experiment attempt ran, so they release the reservation and the same scientific request may be submitted normally after capacity or host inventory recovers. When a peer's compatibility cap is already occupied, another central submission from that peer remains queued until the active process group drains. Other peers may continue to launch, and the blocked job does not consume an attempt or create a capacity-failure record. The final child receives an immutable attempt directory plus exact accelerator variables. Task descendants must preserve them: - `PRAXIST_EXPERIMENT_ID` - `PRAXIST_EXPERIMENT_ATTEMPT_ID` - `PRAXIST_EXPERIMENT_ATTEMPT_DIR` - `PRAXIST_RESOURCE_PROFILE` - `PRAXIST_ASSIGNED_GPU_UUIDS` - `CUDA_VISIBLE_DEVICES` - `NVIDIA_VISIBLE_DEVICES` The built-in Linux observer inventories physical NVIDIA GPU UUIDs. It does not advertise automatic MIG-slice discovery or placement. A task wrapper may remain compatible with an opaque externally supplied identifier, but task initialization must not claim that standard central scheduling validated MIG placement. The submitted evaluator may create ordinary worker or trainer descendants; they inherit the same process group and resource envelope. It must not submit each internal worker as a second top-level experiment. If a compatibility launcher is encountered inside an active attempt, Praxist keeps that child inside the existing allocation instead of recursively queueing it. This reuse is verified against the run-owned attempt directory, committed READY/GO handshake, the caller's live process group, and the scheduler's current in-memory attempt state. Mutable attempt environment variables or copied handshake files alone do not establish ownership. For an active central run, `/resource_scheduler/endpoint.json` is the run-owned launch authority. Peer shell commands may not downgrade that run to legacy launching by changing scheduler environment variables. The environment endpoint remains the compatibility source for legacy callers and runs that do not have run-owned endpoint metadata. ### Optional Managed NVIDIA/CUDA Descendant Binding This subsection applies only after the task's unchanged baseline was observed to use, and task initialization explicitly selected, the compatible Praxist-managed NVIDIA/CUDA backend. `PRAXIST_ASSIGNED_GPU_UUIDS` is then the authoritative ordered physical GPU assignment. A task harness must preserve that exact value in `CUDA_VISIBLE_DEVICES` and `NVIDIA_VISIBLE_DEVICES` across evaluator, trainer, worker, shell, and container boundaries. A framework may use local `cuda:0` inside the mask, but a launcher must not write that local ordinal back into a new child's visibility environment. Missing masks may be restored from the Praxist assignment; conflicting masks must fail clearly instead of silently rebinding. Generated tasks using this backend should carry fast UUID, multi-UUID, missing mask, conflicting mask, standalone, and forced-CPU contract tests. On a host with multiple usable GPUs, launch readiness also includes a bounded non-zero UUID parent/child CUDA check and driver-observed PID-to-UUID comparison. This is a placement-integrity check, not a CPU/accelerator benchmark or training run. CPU-only tasks, unified-memory systems, task-managed devices, and other accelerator backends are valid scheduler paths and do not inherit this UUID contract merely because an accelerator is present. Before the evaluator executes, a small local launch barrier waits until the semantic intent, process group, resource allocation, and protected-process record are durable. Resume rebinds that same allocation before releasing the barrier, so a crash cannot turn one semantic experiment into duplicate work. ## Timing And Close Complete mature evaluations should begin early, not only after assessment reports mature debt. `work-class=mature` has queue priority, while `exploration_reserve` prevents mature work from consuming every slot when exploration is queued. When a configured mature close quorum is still missing at assessment, Praxist stops ordinary queued/new admission but keeps deadline-safe mature top-ups eligible. Assessment is not `CLOSING_SIGNAL`. Once the quota is met, or the generation reaches its safety bound, the normal strict close path takes over. At generation close, Praxist freezes the scheduler queue before writing `CLOSING_SIGNAL`. Queued/new work is rejected; already-running process groups continue to drain and publish evidence. Once protected work is gone, the existing adaptive drain grace bounds agent-only cleanup and passive tool waits; it does not kill evaluator processes. Runtime-private background-task output files are never used as lifecycle facts because empty stdout/stderr is a valid successful result. `praxist stop` similarly freezes all new admission before discovering and terminating scheduler-owned process groups. The existing run shutdown sentinel is the primary fence; a confirmed central-scheduler freeze is an equivalent fallback. If neither succeeds, that run is left untouched and reported in `failed_run_ids` with a nonzero CLI result. Fenced runs receive a bounded rescan until scheduler-owned and exact run-environment descendants are stably absent. A union bulk stop also skips its independent process-name scan when any registry run could not be fenced, because portable hosts may not expose enough evidence to distinguish that run's orphan from an unrelated controller. Explicit process-scan-only operation remains unchanged. A framework-owned `peer_workspaces/` cwd is a narrow fallback for a descendant that cleared its environment; the run root by itself is not process ownership. Process identities are revalidated before every signal, using the portable `ps` start identity when procfs is unavailable, and unrelated operator or monitor processes are not selected by broad command matching. The live central scheduler supplies its complete active process-group set even after a launcher exits. A legacy manifest with a live group but no verifiable launcher is not guessed at or marked stopped; its run is returned in `failed_run_ids` for explicit operator follow-up. ## State And Compatibility Current state is a compact derived view at: ```text /resource_scheduler/status.json ``` `running` is a lifecycle count: the wrapper or a descendant process group is still alive. It is not a GPU-activity count. `running_activity` summarizes the separate observation, and active jobs may include `resource_activity` with `gpu_compute_active`, `gpu_context_idle`, `gpu_context_present`, `no_gpu_process_observed`, `gpu_process_attribution_unavailable`, `non_gpu_allocation`, or `unknown`. These fields are derived telemetry only and never alter maturity, ranking, retries, or result validity. A single `no_gpu_process_observed` sample can be setup, CPU work, evaluation, or a phase transition. `gpu_process_attribution_unavailable` means the accelerator reported a process that could not be mapped into the scheduler's PID namespace; the accompanying `attribution` value distinguishes `complete`, `partial`, and `unavailable` ownership mapping. `unknown` means the underlying observation was unavailable. Admission remains conservative when ownership cannot be mapped so shared-device external load is not mistaken for Praxist work. Use progress, logs, process trees, and result mtimes before diagnosing a stall. On Linux, a process group containing only zombie or exited members is terminal work, even though `killpg(..., 0)` can still report that group as present. Praxist reconciles that kernel state before retaining a running slot, then reaps the wrapper and releases the allocation. A sleeping or otherwise live process is not treated as a zombie, and platforms without procfs retain the portable process-group check. Attempt logs live under `/logs/experiments/`; immutable attempt metadata lives under `/resource_scheduler/attempts/`. Existing protected-PID manifests remain the process-lifecycle compatibility surface used by close, stop, diagnostics, and late/quarantined result handling. The small launch-barrier interpreter runs without Python `site` initialization so a peer's generated `sitecustomize` cannot mistake the trusted READY handshake for a task write. The barrier preserves the submitted environment unchanged when it executes the real task command, so Python task descendants still load the normal runtime guard. Peers never receive write access to scheduler state. Tasks without `mode: central` retain the legacy launch behavior. Central mode does not silently fall back to peer-local launching if its service cannot start. One run-local owner lock prevents two scheduler services from controlling the same queue. Acknowledged submissions and terminal queue rejections are fsynced before their in-memory transitions, so restart replay neither loses accepted work nor revives work rejected by close, freeze, or stop. Runtime environment values that look credential-bearing by name or value shape are stored only as hashes. --- # Scientific Literature and Database Lookup Praxist provides optional public literature, scientific-database, open-access, and provenance lookup through `tool_server:literature_lookup`. It requires no additional API key and does not change the automated experiment loop. ## Capabilities The tool server exposes: - `literature_search(query, sources, max_results)` for normalized cross-source search; - `literature_resolve(identifier)` for DOI, PMID, arXiv, or OpenAlex work identifiers; - `literature_open_access_text(identifier_or_url, max_chars)` for lawful open HTML/XML text or metadata, hash, and provenance for an open PDF; - `scientific_database_search(query, sources, max_results)` for public scientific databases; - `literature_source_guide(domain, objective)` for source-selection and verification guidance. Public sources include arXiv, OpenAlex, PubMed metadata, Crossref, Semantic Scholar metadata, Europe PMC, UniProt, and ClinicalTrials.gov. Individual services may rate-limit or fail. Praxist reports per-source warnings instead of failing an unrelated research run. ## Enable in a Task The standard tool set includes the passive lookup server. A task descriptor should list its complete active tool set so resolve and runtime selection agree: ```yaml praxist_plugins: tools: - tool_server:evaluation_tools - tool_server:frontier_tools - tool_server:finding_graph_query - tool_server:memory_tools - tool_server:prior_work_tools - tool_server:run_report - tool_server:literature_lookup ``` The tool is passive: network access occurs only when an agent explicitly calls it. A task may define a task-local literature role, but Praxist core must not contain domain-specific search strategy. Principal Investigator (PI) memo agents and peers with the tool perform lookup; the Chair planning agent synthesizes the evidence provided to it. ## Current-Environment-Only Rule Search results may mention datasets, checkpoints, simulators, dependencies, licenses, APIs, or environments that are not available locally. During a run, agents must not acquire or install those missing resources. They should extract the useful scientific idea and adapt it to the task's existing data, evaluator, dependencies, hardware, and runtime. Missing resources may be recorded as task-local limitations or future requirements. They are not permission to mutate the host. ## Evidence and Provenance Literature is context, not measured task performance. Records should preserve: - title, authors, year, venue, and stable identifier; - source and retrieval time; - open-access status; - exact claim supported by the source; - uncertainty, contradiction, or negative evidence; - a pointer that lets a later agent retrieve the original record. A negative lookup result is also evidence. Report the query, sources attempted, warnings, and coverage limits. Do not rewrite "not found in searched sources" as "does not exist." ## Runtime Behavior The lookup server: - uses bounded timeouts and result limits; - normalizes records without claiming cross-source identity when uncertain; - degrades one failing source without disabling other tools; - does not bypass paywalls or authentication controls; - does not download task data or install dependencies; - keeps retrieved material separate from evaluator evidence. Use `praxist-scientific-research` when an operator agent should gather task context before a run. Use the tool server when a running peer or PI needs a focused source check. --- # Troubleshooting Use the smallest command that can identify the failing boundary. Do not edit run artifacts to make a failed check appear healthy. ## Host or Authentication Is Not Ready ```bash praxist doctor --json ``` For Codex-native mode: ```bash praxist setup --profile codex-native --install-skills codex # or: claude praxist doctor --codex-native --task-path /absolute/path/to/task --json ``` If a tested runtime package is missing or mismatched, invoke the `praxist-runtime-install` skill instead of independently upgrading an SDK. Codex-native authentication behavior and diagnostic-override semantics are defined in [Credentials](https://praxist.sapient.inc/en/docs/guides/credentials#codex-native-mode-authentication). Do not apply that repair to another setup profile. ## Package Download Certificate Failure If installation reports `SSLCertVerificationError` or `CERTIFICATE_VERIFY_FAILED`, repair the selected Python trust store and rerun the pip command. On macOS, python.org distributions provide an `Install Certificates.command` alongside the installed Python. Managed Python or corporate environments should use their supported CA-bundle configuration. Do not bypass TLS verification with `trusted-host`, and never place an API key in a command argument while troubleshooting connectivity. ## Task Does Not Resolve ```bash praxist resolve /absolute/path/to/task ``` Resolution makes no LLM calls. It reports invalid task configuration, missing plugin descriptors, unresolved task-local references, and unsupported agent runtime/API provider combinations before launch. ## A Started Run Disappears `praxist start --daemonize --json` reports process creation before every research stage necessarily initializes. Inspect: ```bash praxist status --json praxist --monitor --latest ``` Then read the run's `run_summary.json` and launcher log path reported by start. A stale registry record means the registered process is no longer alive; it is not proof that the research completed. ## Run Appears Stalled Use: Invoke the `praxist-diagnostic` skill in the current agent. The diagnostic workflow separates a long active experiment from missing stage artifacts, resource starvation, runtime/API friction, blocked generation close, or incomplete evidence. The foreground monitor is observational and must not be used as scientific evidence. ## Stop or Resume Prefer the lifecycle skill: Ask the `praxist-control` skill to stop the current run or resume the latest run. Interrupted generation, Principal Investigator (PI) panel, and Gems boundaries require artifact-aware inspection before `resume`. The canonical procedure is defined in [Direct CLI Operations](https://praxist.sapient.inc/en/docs/guides/operators). ## Documentation Build Fails ```bash uv sync --extra docs uv run python scripts/build_docs_site.py ``` The build is strict. It fails on stale generated references, unowned pages, duplicate navigation ownership, broken local links, or MkDocs warnings. For configuration and credential details, use [Configuration Discipline](https://praxist.sapient.inc/en/docs/concepts/config_discipline) and [Credentials](https://praxist.sapient.inc/en/docs/guides/credentials). --- # Platform Support ## Release Validation Matrix | Operator host | Release status | Required verification | |---|---|---| | Linux on CPython 3.11 or 3.12 | Continuously tested by release CI | Run `praxist doctor` for runtime and provider readiness | | macOS on CPython 3.11+ | Package and CLI compatibility target; not continuously tested by release CI | Run `praxist doctor` and validate all task-owned dependencies | | Linux on other CPython 3.11+ versions | Package compatibility target; not continuously tested by release CI | Run `praxist doctor` and a task-specific smoke test | | Windows-native | Outside the current research-runtime contract | Use a supported Linux environment instead | The package metadata accepts CPython 3.11 or newer, but that compatibility range must not be read as a claim that every interpreter, operating-system and hardware combination has passed release CI. Codex or Claude Code must already be installed and usable for skill-driven operation. Headless Linux and remote shells are normal environments. ## Research Hardware Praxist does not require a particular accelerator. CPU-only systems, macOS unified memory, NVIDIA/CUDA, task-managed accelerators, and other task-owned backends are valid when the research project itself supports them. Task initialization observes the unchanged baseline on the current host before declaring resource behavior. It must not infer a platform from product names, compare artificial CPU-only and accelerator-only rewrites, or invent an accelerator profile. The central scheduler manages only resource classes explicitly represented by the task harness. Its full contract is documented in [Central Experiment Scheduler](https://praxist.sapient.inc/en/docs/guides/central-resource-scheduler). ## Out of Scope The Praxist package and setup wizard do not provide: - GPU drivers or CUDA; - model-training frameworks; - datasets or simulators; - task-specific containers; - cluster schedulers. Those belong to the research project or host administrator. `praxist resolve` can still parse task and plugin configuration without loading POSIX locking at module import time. `praxist doctor` reports unsupported native platforms before a research launch, and an attempted central-scheduler run fails with a direct platform message rather than silently weakening host locking. ## Filesystem and Process Expectations Praxist supports ordinary local paths and symlinked task or experiment storage. The operator must have permission to create task run directories, user configuration and registry state. The selected Python environment must be writable by its owner. Daemonized runs are independent of the agent conversation that launched them. The live monitor is a separate foreground process and `Ctrl-C` exits only that monitor. --- # Product Usage Controls Praxist includes optional, pseudonymized product-usage reporting. Collection is off until the current operating-system user explicitly consents, and declining or withdrawing does not affect installation or research. Accepting the [Praxist User Agreement](https://praxist.sapient.inc/en/docs/legal/user-agreement) or the [Fair Source License](https://github.com/sapientinc/praxist/blob/main/LICENSE.md) does not enable collection. Read the [Privacy Notice](https://praxist.sapient.inc/en/docs/legal/PRIVACY) for the complete privacy policy, purposes, retention rules, and user rights. The [Praxist User Data Collection Notice](https://praxist.sapient.inc/en/docs/legal/product-usage-data-notice) is the exact versioned text presented when consent is requested. The [Product Usage Technical Documentation](https://praxist.sapient.inc/en/docs/operations/DOCUMENTATION) describes the implementation, endpoint, schema, storage, and failure isolation. ## Review and Choose Review the current in-product notice and record a choice: ```bash praxist product-usage notice praxist product-usage consent ``` Inspect or withdraw the choice at any time: ```bash praxist product-usage status --json praxist product-usage withdraw ``` A plain pip installation and every non-interactive path leave consent unset. Interactive setup presents the notice in a scrollable local terminal view. Withdrawal stops future capture and removes unsent local events. See the Privacy Notice for the treatment of events already delivered. ## Collector Development Install server dependencies only on a collector development or deployment host: ```bash uv sync --group dev --extra product-usage-server export DATABASE_URL='postgresql+psycopg://user:password@host/database' export COLLECTOR_INGESTION_ENABLED=true export COLLECTOR_MAX_TABLE_BYTES=$((2 * 1024 * 1024 * 1024)) uv run alembic -c services/product_usage/alembic.ini upgrade head uv run praxist-collector ``` Run retention in a separate process or container: ```bash uv run praxist-retention ``` Collector ingestion is disabled until explicitly enabled. Deployment assets and their operational controls are under `services/product_usage/`; implementation details and audit entry points are owned by the technical documentation. --- # Praxist Data-Collection Module: Technical Documentation **Last updated:** 27 August 2026 **Purpose:** As required by Section 1.4 (Data Collection) of the Fair Source License Agreement (Version 1.0), this document describes the source-code structure of the data-collection module and discloses the data-receiving endpoint URL. It is written for users, enterprise IT, and security auditors. **Related documents:** [Privacy Notice](https://praxist.sapient.inc/en/docs/legal/PRIVACY); [Praxist User Data Collection Notice](https://praxist.sapient.inc/en/docs/legal/product-usage-data-notice) (in-product notice text, Notice version 3) > This document describes the current implementation. Keep it and the Privacy Notice aligned with changes to the product-usage contract. --- ## 1. Module Map | Part | Path | Notes | | --- | --- | --- | | Client SDK | `praxist/product_usage/` | Consent management, environment identity, event generation, local outbound queue, upload | | Client integration | `praxist/infrastructure/product_usage.py` | Observer that projects Research Run lifecycle into telemetry events; failures are isolated from the Research Run | | CLI entry point | `praxist/cli/product_usage.py` | The `praxist product-usage` command family (notice / consent / status / withdraw) | | Server Collector | `praxist/product_usage/app.py`, `collector.py`, `postgres.py`, `retention.py` | HTTP ingestion, schema validation, idempotent persistence, retention deletion | | Server deployment | `services/product_usage/` | Dockerfile, Nginx configuration, Compose, deployment scripts | | Protocol & schema | `praxist/product_usage/protocol.py`, `schemas/v2/*.json` | Closed Schema V2 (shared by client and server) | | Legal text | `docs/legal/product-usage-data-notice.md` | The notice text shown in-product (Notice version 3) | ## 2. Client File Map | File | Responsibility | | --- | --- | | `consent.py` | Consent-state storage: a `unset` / `granted` / `denied` state machine that fails closed; atomic writes (0600 permissions); consent records are bound to the Notice version — a version mismatch counts as no consent; Agent-assisted replies recognize only `Yes` / `Agree` / `No` / `Disagree` | | `identity.py` | Environment identity: generated at random via UUIDv4 and persisted locally in `environment.json`; never derived from any personal, device, or task information | | `paths.py` | Fixed per-OS local file paths (Section 5); no environment-variable or project-level overrides | | `lifecycle.py` | Generation of run-level event IDs, telemetry run IDs, event sequence numbers, and the four lifecycle events | | `outbox.py` | Bounded local SQLite outbound queue (offline buffering, later delivery, cleared on withdrawal) | | `batching.py` | Bounded JSON batch encoding (parsed identically on the server) | | `transport.py` | Endpoint selection and HTTP sending (Section 3) | | `client.py` | The `UsageSdk` facade: every collection failure is isolated and never reaches the Research Run | | `protocol.py` | Closed Schema V2 models and boundary constants (Section 4) | | `notice.py` | Loads the in-product notice text | | `app.py` / `collector.py` / `postgres.py` / `retention.py` | **Server side**: HTTP entry, validation and idempotency core, PostgreSQL persistence, retention-deletion job | ## 3. Data-Receiving Endpoint URLs | Environment | Endpoint | Notes | | --- | --- | --- | | **Production (release builds)** | `https://telemetry.theaiscientist.com/v1/events` | HTTPS encryption, server certificate verification, no redirect following | | **Development (internal `.dev` builds only)** | Internal development collector (address not published) | Plain HTTP; handles only internal development/test data, never user data; not shipped with release builds | Endpoint selection: `default_batch_sender()` in `transport.py` chooses the sender based on whether the Praxist version string contains `.dev`. The production endpoint must be a valid HTTPS URL or construction is refused outright. Request contract: - `POST` with `Content-Type: application/json`; request body capped at 32 KB; at most 50 events per batch; - carries only the fixed, protocol-level User-Agent `Praxist-Product-Usage/2` — no cookies and no additional request headers; - network timeout is 2 seconds; a success response is `202` with body `{"accepted": n, "duplicates": n}`; - error responses: `400` (malformed request / unsupported schema version), `413` (too large), `415` (non-JSON), `503` (ingestion paused or temporarily unavailable); - transmission failures never block or affect the Research Run; undelivered events stay in the local outbound queue and are sent automatically once connectivity returns. ## 4. Closed Schema V2 and Boundary Constants The schema is a closed model (`extra="forbid"`): both client and server reject any out-of-schema field, and schema extensions require an explicit change to `protocol.py` plus a version bump. | Constant | Value | Meaning | | --- | --- | --- | | `SCHEMA_VERSION` | 2 | Event-structure version | | `CONSENT_NOTICE_VERSION` | 3 | Current Notice version (consent records are bound to it) | | `MAX_BATCH_EVENTS` | 50 | Maximum events per batch | | `MAX_REQUEST_BYTES` | 32 KB | Maximum request-body size | | `MAX_ERROR_SUMMARIES` | 16 | Maximum error-summary groups per event | | `MAX_ERROR_COUNT` | 65,535 | Maximum count per error group | | `MAX_DURATION_MINUTES` | 43,200 (30 days) | Maximum recorded active-run duration | The closed value lists for event fields and error categories are in Section 3 of the [Privacy Notice](https://praxist.sapient.inc/en/docs/legal/PRIVACY) and in `schemas/v2/usage-event.schema.json` and `schemas/v2/usage-batch.schema.json`. ## 5. Consent State and Local File Paths All paths below are fixed, current-user-only locations with 0600/0700 permissions: | Platform | Consent record | Outbound queue / environment identity | | --- | --- | --- | | Linux | `~/.config/praxist/product-usage/consent.json` | `~/.local/share/praxist/product-usage/outbox.sqlite3`, `environment.json` | | macOS | `~/Library/Application Support/Praxist/product-usage/consent.json` | `outbox.sqlite3`, `environment.json` in the same directory | | Windows | `%LOCALAPPDATA%\Praxist\product-usage\consent.json` | `outbox.sqlite3`, `environment.json` in the same directory | - **Run-state files:** `runs/.json` alongside `environment.json`; stored locally only and never uploaded. **How the hash is computed** (see `run_state_path()` in `praxist/product_usage/paths.py`): the run-directory path is first normalized with `expanduser` and `resolve` (expanding the user directory and resolving it to a canonical absolute path), then UTF-8 encoded and hashed with SHA-256; the 64-character lowercase hexadecimal digest becomes the filename. SHA-256 is one-way, so the original path cannot be recovered from the filename, and since the file never leaves the machine, no path information is exposed; - **CLI:** `praxist product-usage notice | consent | status --json | withdraw`; - **First use:** the notice is shown and an explicit choice is awaited only in an interactive terminal while the state is `unset`; non-interactive environments remain `unset` (i.e., nothing is collected). ## 6. Offline and Failure Behavior - **Offline or upload failure:** events are buffered in the local outbound queue and delivered when connectivity returns — **research functionality is entirely unaffected**; - **Withdrawal (`withdraw`):** immediately stops all future capture and deletes every local unsent event; - All client-side collection/upload exceptions are isolated by the fail-closed facade in `client.py` and never interrupt a Research Run. ## 7. Server-Side Privacy Measures - Nginx reverse proxy: `access_log off`; strips the `X-Forwarded-For`, `X-Real-IP`, `Cookie`, and `User-Agent` headers before requests reach the application; - Two-level rate limiting (per client and global); storage capacity ceiling `COLLECTOR_MAX_TABLE_BYTES` (default 2 GB) — new events are refused beyond it; - Master ingestion switch `COLLECTOR_INGESTION_ENABLED`, which can pause ingestion entirely (returns 503); - The Collector container binds to `127.0.0.1` only and reaches the managed PostgreSQL over a private endpoint; the database is never exposed to the public internet; - `received_at` is generated by the server after validation and cannot be supplied or altered by the client; - Retention job: runs once at service start and at least every 24 hours thereafter; events enter the deletion window on day 179 (one day of scheduling slack against the stated 180-day ceiling); if the retention job fails, it exits and the container restarts to retry; - The server has **no** interface for deleting already-delivered events by environment identifier (a design constraint disclosed in Section 8 of the [Privacy Notice](https://praxist.sapient.inc/en/docs/legal/PRIVACY)). ## 8. How to Audit It Yourself 1. See the exact fields that would be reported: `praxist/product_usage/schemas/v2/*.json` and `protocol.py`; 2. Check current consent status: `praxist product-usage status --json`; 3. Inspect local files: the paths listed above (all readable and writable only by the current user); 4. Packet-capture verification: the endpoint, the fixed User-Agent, and the request-body content can all be verified with standard network capture tooling; 5. The complete notice text: `praxist product-usage notice`, or `docs/legal/product-usage-data-notice.md` in the repository. ## 9. Versioning and Change Management - The schema version, Notice version, and all boundary constants are defined centrally in `praxist/product_usage/protocol.py`; - When the notice content changes, `CONSENT_NOTICE_VERSION` increases; previously recorded consent does not carry over to the new version and is requested again at first use; - If the implementation described in this document changes, this document, the [Privacy Notice](https://praxist.sapient.inc/en/docs/legal/PRIVACY), and the in-product notice must be updated together. --- **Contact:** praxist@sapient.inc --- # Agent Runtimes Agent runtimes execute `AgentRunRequest`, emit normalized `AgentEvent` records, and return `AgentRunResult`. The runtime controls the agent session; the API provider (`model_provider:*`) separately describes API shape, endpoint, model defaults, and credentials. ## Runtime Contract Every production runtime adapter is responsible for: - translating Praxist prompt, model, tool, cache, sandbox, and timeout intent; - exposing selected MCP tool servers without leaking API provider response objects; - normalizing assistant text, tool calls/results, errors, usage, and terminal state; - preserving cancellation and timeout status; - redacting credentials before events enter trajectory or logs; - supporting concurrent long-running peer sessions within its declared capacity. Exact capabilities vary by SDK. A runtime must report an unsupported contract or unknown usage explicitly rather than pretending the capability exists. ## Bundled Runtimes - `agent_runtime:claude_sdk` is the default and recommended production runtime for new task projects, tested with `claude-agent-sdk==0.2.136`. - `agent_runtime:codex_sdk` is an explicitly selected production runtime built on the official `openai-codex==0.147.0` Python SDK. - `agent_runtime:fake_runtime` is the deterministic offline runtime used by conformance tests. Selecting `codex_sdk` does not change the default runtime for existing tasks. ## Claude SDK Liveness Each `agent_runtime:claude_sdk` session consumes its SDK stream on an isolated worker loop while the research loop retains timeout and cancellation authority. The adapter tracks complete SDK messages, partial model-stream events, foreground tool activity, and protected background work as distinct progress signals. Partial events are observability-only and are not copied into the canonical agent transcript. A liveness warning is emitted only when the observed state is `model_waiting` and every progress source has remained silent past the warning interval. A long-running foreground tool or active background task is reported as a low-frequency healthy-work state instead of an SDK stream stall. Tool names may appear in that health record, but tool inputs, commands, task IDs, and other sensitive payloads do not. These observations do not extend or replace the configured runtime timeout, generation deadline, stop request, or cancellation path. ## Codex SDK Architecture `agent_runtime:codex_sdk` uses long-lived local Codex app-server clients rather than launching a fresh CLI command for every peer turn. An agent runtime/API provider/credential scope may share one client while each request receives an independent ephemeral Codex thread. The adapter consumes typed app-server notifications and maps them to Praxist events. Selected Praxist tool servers are attached directly as stdio MCP servers. Tool allow/deny metadata is translated into the app-server configuration; no shell bridge is part of the runtime contract. The runtime also provides: - streaming assistant, tool, reasoning, plan, file-change, usage, and terminal notifications when the SDK emits them; - turn interruption for Praxist stop requests and timeouts, followed by a bounded drain; - replacement of an unhealthy app-server client without invalidating healthy concurrent turns prematurely; - a private worker pool and bounded stream concurrency so large peer cohorts do not starve unrelated Praxist async work; - runtime-scoped state under `/runtime_state/codex_sdk/`; - native OpenAI authentication through either `OPENAI_API_KEY` or a saved ChatGPT login owned by the SDK-bundled Codex binary; - a read-only account model-catalog probe so Codex-native mode launchers can reject unsupported explicit models before starting a peer cohort; - Praxist sandbox-intent translation for read-only, workspace-write, and full access modes. Codex turns keep strict caller `output_schema` values. When a known non-strict object schema omits `additionalProperties: false`, Praxist leaves it out of the first endpoint request instead of making a predictably rejected call. Existing prompt parsing and task validation remain in force, and one runtime warning records the compatibility fallback. Unrelated API provider errors and strict-schema errors are not retried or hidden. Tasks should not shadow the bundled runtime with model-specific schema adapters. Usage is exact only when the app-server publishes token-usage notifications. Otherwise the normal Praxist `usage_unknown` behavior applies. Prompt-cache behavior remains agent runtime/API provider managed. Codex has a built-in shell surface, so requests that require a provably shell-free runtime are rejected instead of being represented as equivalent to Claude SDK behavior. For Codex-native mode runs, Praxist automatically uses lossless finding-event batching between independent threads. It keeps the complete task contract and canonical artifacts, supplies a larger bounded set of unseen finding references, and asks continuation sessions to reopen exact originals only when needed. This reduces repeated thread bootstrap work without retaining an indefinitely growing Codex conversation or relying on lossy compaction. Stop, closing, resource-supply, timeout, and recovery semantics are unchanged. ## Capability Alignment And Differences Both production adapters target the same Praxist request/result contract, but they are not interchangeable implementations: | Capability | `claude_sdk` | `codex_sdk` | | --- | --- | --- | | Selected Praxist MCP tools | Direct SDK MCP integration | Direct app-server stdio MCP integration | | Streaming | Normalized from Claude SDK messages | Normalized from typed app-server notifications | | Timeout and stop | Runtime-specific cancellation path | Turn interrupt plus bounded notification drain | | Usage | Recorded when the SDK/API provider exposes it; otherwise unknown | Token-usage notifications when present; otherwise unknown | | Sandbox intent | Claude SDK permission/sandbox integration | Codex read-only/workspace/full mapping; built-in shell remains part of the runtime | | Long-run concurrency | Independent concurrent peer sessions | Independent threads over shared long-lived clients with bounded stream concurrency | | Non-native API providers | Claude SDK/API provider compatibility path | Responses-to-Chat relay for supported Chat Completions API providers | Do not claim complete behavioral equivalence. Tool naming, event granularity, sandbox capabilities, cache behavior, API provider errors, and usage availability remain SDK-specific even though Praxist normalizes their durable result shape. ## API Provider Routing The Codex app-server consumes the Responses protocol. API provider routing is: | API provider shape | Codex SDK path | | --- | --- | | OpenAI / `model_provider:openai_compatible` | Direct SDK/app-server connection | | `model_provider:deepseek_alias` | Private run-scoped `codex-relay` to DeepSeek Chat Completions | | `model_provider:openrouter` | Private run-scoped `codex-relay` to OpenRouter Chat Completions | Praxist starts and stops the relay; operators must not launch a relay per peer. The relay listens only on an ephemeral local port and receives only the selected API provider credential. An API provider not declared compatible by the runtime plugin must fail during resolution or startup rather than being silently rerouted. For OpenRouter only, the relay adds a hashed run-scoped `session_id` for sticky routing and cache locality. It does not enable response caching. DeepSeek relay reasoning overrides are added only when the task selects an explicit policy; `auto` preserves the existing route behavior. Lossless context-efficiency controls are documented in [Cost Optimization](https://praxist.sapient.inc/en/docs/guides/cost-optimization). The automatic policy applies to Codex-native mode and OpenRouter routes; direct DeepSeek is explicitly excluded. ## Reasoning Policy Task projects may set one reasoning policy across API providers for every peer, Principal Investigator (PI), Chair, and Deep Innovation Gate (DIG) planner call: ```yaml agent: reasoning_effort: max # auto | off | low | high | max ``` `max` is the default for new and existing task projects that omit the field. It requests the strongest reasoning level supported by the selected route. `auto` is an explicit opt-in to the API provider/agent runtime's native default. `off` explicitly disables model reasoning when the route supports that control; `low` and `high` request the corresponding effort. The legacy `premium_mode: true` setting remains supported as `max` when `reasoning_effort` is `auto`; an explicit non-`auto` policy takes precedence. For `claude_sdk` with DeepSeek, Praxist maps this policy to DeepSeek's Anthropic compatibility fields: `thinking.type` is `enabled` or `disabled`, and enabled requests carry the selected effort. Other Claude-compatible API providers retain their native adaptive-thinking mapping. For `codex_sdk` with DeepSeek, the private run-scoped relay injects the same policy into Chat Completions requests and retains the API provider's reasoning state across tool-call subrequests. Praxist does not summarize or reconstruct that state. Native Codex models, including Codex-native `gpt-5.6-luna`, receive the closest supported SDK effort (`max` maps to `xhigh`). OpenRouter's relay route uses its unified `reasoning.effort` object. Models whose API provider contract requires reasoning may reject `off`; Praxist surfaces that API provider error and leaves the user's policy unchanged. Reasoning controls change model behavior and may change latency, output-token use, and cost. They do not change task evidence, promotion, timeout, or generation-close contracts. ## Install And Select Install the Codex runtime extra in the Praxist environment: ```bash python -m pip install 'praxist[codex]' ``` The extra pins `openai-codex==0.147.0`, `claude-agent-sdk==0.2.136`, and `codex-relay==0.5.5`, and includes the MCP dependencies used by bundled Praxist tools. These versions are the tested runtime compatibility baseline; upgrade them only together with Praxist runtime validation. The extra does not install a separate task environment. Select a runtime through a setup profile, task configuration, or explicit start override. [Credentials](https://praxist.sapient.inc/en/docs/guides/credentials) owns API-key and saved-login setup, authentication precedence, private Codex homes, and verification. The [Quickstart](https://praxist.sapient.inc/en/docs/getting-started/quickstart) owns first-use profile selection; the generated [CLI Reference](https://praxist.sapient.inc/en/docs/reference/cli) owns exact command options. ## Direct Agent Skill Use The human-facing Codex or Claude Code CLI can invoke the bundled Praxist skills directly. This operator host is independent of the peer runtime. A saved Codex ChatGPT login may authenticate the official Codex SDK runtime when native OpenAI is explicitly selected, even when Claude Code hosts the operator workflow. Peer execution still happens through Praxist-owned runtime clients, not by attaching to the interactive operator session. ## Adapter Checklist When adding or changing a runtime adapter: 1. accept the runtime-neutral request/context contract; 2. translate model, prompt, MCP, sandbox, cache, timeout, and tool options; 3. emit normalized typed events and a normalized terminal result; 4. record usage when available and unknown usage otherwise; 5. redact secrets and API provider response objects; 6. preserve cancellation, timeout, and concurrent-turn isolation; 7. add offline conformance plus focused API provider/MCP integration coverage; 8. document capability differences instead of claiming cross-SDK equivalence. --- # Open-Source Model APIs For sustained research, Praxist generally favors APIs serving capable open-source or open-weight models when a representative run demonstrates high cache reuse, sufficient research quality, and stable throughput. This is a selection policy, not a hidden runtime default: the operator still chooses the profile during setup. ## Current Shortlist | Priority | Option | Selection note | |---|---|---| | 1 | **DeepSeek V4 Pro** | Praxist provides a maintained direct API profile. Evaluate it first where the service is available and appropriate for the project. | | 2 | **Open-source models through OpenRouter** | Use the OpenRouter profile when routing flexibility matters. Select the exact model explicitly and verify that its route reports useful cache reuse. | | 3 | **Operator-managed open-source model endpoints** | Add or select a compatible provider plugin when deployment policy requires a private or self-hosted endpoint. Validate the plugin contract before a long run. | The order is a practical starting point, not a claim that one model is best for every task. Availability, pricing, model behavior, and provider-side caching can change independently. ## Validate Before a Long Run 1. Run a short, representative workload through the intended agent runtime and API route. 2. Confirm output quality against the task's normal evaluator rather than a provider-specific proxy. 3. Inspect cached and uncached input usage, latency, and failures in the run artifacts. 4. Keep the selected route only when the total cost and research quality are acceptable together. Praxist preserves stable prompt prefixes where the selected route supports them, but no model name alone guarantees a high cache-hit rate. See [Cost Optimization](https://praxist.sapient.inc/en/docs/guides/cost-optimization) for the cache contract and [API Providers](https://praxist.sapient.inc/en/docs/guides/model-providers) for supported provider shapes. --- # API Providers `model_provider:*` API provider plugins describe API shape, model defaults, credential requirements, cache capability, and route-specific compatibility. Agent runtime plugins execute the agent loops. ## Built-In API Provider Shapes - `model_provider:openrouter` for OpenRouter-routed model names. - `model_provider:openai_compatible` for OpenAI-compatible endpoints. - `model_provider:anthropic_messages` for native Anthropic Messages style. - `model_provider:deepseek_alias` for DeepSeek-compatible aliases. API provider names represent API format and routing. A task or operator may override the `ModelProfile` used by a stage. ## API Provider Manifest Expectations An API provider manifest should declare: - supported API format; - default model, if any; - endpoint base, if fixed; - required credential refs; - cache capability; - usage reporting capability; - compatibility with agent runtime plugins. ## Multi-Model Runs Research-loop agents may use different model profiles when the task contract and selected runtime support them. Peer exploration and planning roles should resolve providers and model names through the same configuration boundary. Do not hard-code a task-specific model inside a generic API provider plugin. Run-wide reasoning effort belongs to the agent runtime policy. Configure it under `agent.reasoning_effort` as documented in [Agent Runtimes](https://praxist.sapient.inc/en/docs/guides/agent-runtimes#reasoning-policy); adapters translate that single policy to each API provider's supported wire contract. The default is `max`; select `auto` explicitly to retain an API provider's native effort default. ## Provider Conformance API provider tests should cover: - credential resolution and redaction; - API provider/agent runtime compatibility; - cache capability mapping; - missing or invalid key diagnostics; - usage unknown behavior when the API provider does not return metering. --- # Credentials Credentials are resolved by Python startup code and represented by redacted credential references. Shell wrappers do not own credential behavior. ## API Provider Credentials Set an API key for the selected API provider before starting a run. The maintained direct DeepSeek V4 Pro route uses `model_provider:deepseek_alias` with `agent_runtime:claude_sdk`: ```bash export DEEPSEEK_API_KEY=... praxist start --model-provider model_provider:deepseek_alias \ --runtime agent_runtime:claude_sdk \ --model deepseek-v4-pro ``` See [Open-Source Model APIs](https://praxist.sapient.inc/en/docs/guides/open-source-model-apis) for route-selection criteria. Credential handling is identical regardless of which API-backed profile the operator selects. When an operator explicitly selects `agent_runtime:codex_sdk`, the same `DEEPSEEK_API_KEY` is scoped to a private run-local `codex-relay` because DeepSeek exposes Chat Completions and the Codex app-server expects Responses. OpenAI uses `OPENAI_API_KEY` directly without the relay. Neither path stores a raw key in task files or human Codex CLI sessions. ### Codex-native mode authentication Codex-native mode uses `agent_runtime:codex_sdk` with a saved ChatGPT login for `model_provider:openai_compatible` when `OPENAI_API_KEY` is absent: ```bash praxist setup --profile codex-native --install-skills codex praxist start \ --codex-native \ --agent-system codex_sdk \ --model-provider model_provider:openai_compatible \ --task-path /path/to/task-project ``` The setup command verifies the Codex binary pinned inside the installed SDK, not an unrelated executable found first on `PATH`. If that binary currently uses API-key authentication, a local interactive terminal opens its login flow with API provider key environment variables removed, then verifies that ChatGPT authentication is active. A valid existing ChatGPT login is reused without a new prompt. Noninteractive environments report the required local setup command instead of partially configuring the profile. This subscription check is exclusive to an explicitly selected Codex-native profile or `--codex-native` operation. Ordinary `codex_sdk` API provider routes do not fail readiness because of the operator's Codex login method. `--codex-native` is authoritative after user and task configuration files are loaded: inherited API provider/agent runtime/model defaults and API-key/custom-endpoint variables cannot silently switch or misconfigure this run. An explicit CLI `--model` remains authoritative. Outside this explicit mode, environment configuration and credentials retain their normal precedence. Praxist records only a redacted identity; for file-based login it includes a hash of the stable account identifier, never a token. At runtime Praxist stages `auth.json`, when present, inside a private disposable OS-temporary Codex home so the app-server can refresh its own copy without writing the operator's `CODEX_HOME`. Keyring-backed login uses the same private empty home and the operating-system credential store. The private home is removed when its app-server closes and is never put in task files, run artifacts, logs, or replay. The runtime verifies the app-server account is actually `chatgpt` before starting a peer turn, blanks API-key endpoint overrides, and never falls back to an API or relay when Codex-native mode was selected. On resume, `--codex-native` may select saved-login authentication only when the existing run already has the canonical `codex_sdk` agent runtime and native OpenAI API provider. Resume never rewrites a historical run's agent runtime or API provider; start a new run to change either canonical choice. Other API providers remain supported when the operator explicitly chooses them: ```bash export OPENROUTER_API_KEY=... export ANTHROPIC_API_KEY=... ``` Praxist reads `${XDG_CONFIG_HOME:-$HOME/.config}/praxist/env` by default. Use `PRAXIST_CONFIG_FILE=/path/to/env` or command-local `--config-file /path/to/env` for another configuration. The command-local flag takes precedence; exported process credentials take precedence over file values. `praxist configure-llm` manages built-in API provider configurations only. Custom `model_provider` plugins remain supported through each plugin's documented task or host environment contract; Praxist does not infer custom key-variable names. Do not commit keys, paste keys into logs, or write keys into task files. ## Credential Failover Boundary Single-key quickstart is supported. Built-in environment discovery currently loads at most one credential per API provider. `CredentialFailoverManager` can select a fallback when its caller supplies multiple credentials with matching scope and API provider; an unset `target_ref` acts as a target wildcard. The caller must also record a supported failure. Automatic runtime failure-triggered fallback remains disabled, so this is not a user-selectable runtime mode. Selection and failure state use redacted `CredentialRef` values. ## Tool Credentials Tool-scoped keys are separate from API provider keys. The bundled `tool_server:literature_lookup` is no-key-first: it uses public endpoints such as arXiv, OpenAlex, PubMed metadata, and Crossref-style DOI metadata without requiring task authors to configure another API provider key. Future task-local or external plugins may add service-specific credentials for higher rate limits or licensed sources. Those credentials are optional enhancers, not generic Praxist requirements. Missing tool credentials must disable or degrade only the affected lookup path and must not break tasks that do not explicitly require that source. ## Redaction Trajectory, logs, docs, generated sites, replay reports, and task templates must not contain raw secrets. Tests under `tests/hardening` enforce this boundary. --- # Cost Optimization This guide describes low-risk token optimizations that preserve research facts while reducing repeated context inflation. ## Goals Praxist cost optimization follows the result-preservation principle: - keep peer outputs and raw evidence available; - return compact summaries by default; - make full data available through explicit lookup; - avoid task-specific logic in core; - avoid asking agents to re-read large logs or broad JSON files when a small summary is enough. ## Lossless Session Efficiency By API Provider Praxist automatically enables lossless session efficiency for two expensive routes: - Codex-native mode (`agent_runtime:codex_sdk` with native OpenAI and a saved ChatGPT login); - `model_provider:openrouter` on any supported runtime. Direct `model_provider:deepseek_alias` runs are always excluded. Their event cadence, memory limits, and prompts remain unchanged, even if an operator sets the lossless override. The policy does not compress history or remove findings. It changes how peers consume the same canonical artifacts: 1. stop, closing, and resource-supply events remain immediate; 2. individual shared-finding events within a short interval are collected and followed by one continuation session; 3. the next prompt carries a larger bounded batch of unseen finding IDs plus the existing peer state and handoff; 4. the complete task prompt remains present; 5. the continuation is told to use exact references first and reopen original artifacts whenever details are needed or uncertain. An already-consumed finding suppresses a duplicate wake only when both its explicit identity and content version are unchanged. A corrected or expanded payload with the same identity is surfaced again; missing or unparseable identity fails open and wakes the peer. The canonical findings, full tool outputs, session logs, and task documents remain on disk. This is event coalescing and reference-first navigation, not context compression. Configuration: ```bash # Default: auto-detect Codex-native mode and OpenRouter routes. export PRAXIST_CONTEXT_EFFICIENCY_MODE=auto # Change the finding-only batching interval (default 300 seconds). export PRAXIST_CONTEXT_EFFICIENCY_MIN_SESSION_INTERVAL_SECONDS=300 # Disable finding batching for a comparison run. export PRAXIST_CONTEXT_EFFICIENCY_MODE=off ``` `lossless` explicitly enables the policy for a non-DeepSeek route. Unknown mode values fall back to `auto` rather than blocking a run. When OpenRouter is selected with `agent_runtime:codex_sdk`, its existing private relay additionally receives a non-secret, run-scoped `session_id`. This provides sticky routing for prompt-cache locality. Praxist does not enable OpenRouter response caching: research replies and tool calls must not be replayed verbatim. See the [OpenRouter prompt-caching guide](https://openrouter.ai/docs/guides/best-practices/prompt-caching) and [response-caching distinction](https://openrouter.ai/docs/guides/features/response-caching). Native OpenAI prompt caching remains automatic. Cache hits require exact prefix matches, so stable instructions should precede dynamic generation/session content. See the [OpenAI prompt-caching guide](https://developers.openai.com/api/docs/guides/prompt-caching). ## Tool Output Limits And Full Lookup Problem: MCP tools such as leaderboard, frontier, and finding-graph queries can return large JSON payloads. Even when an agent only needs the top few records, the whole response is fed back into the LLM context. Design: - tool handlers keep their existing business fields, such as `entries`, `pareto_front`, `neighbor_findings`, `nodes`, and `edges`; - handlers attach `_tool_output` metadata with: - `schema_version`; - `tool_name`; - `view = summary`; - `truncated` and `truncated_lists`; - `full_result_ref`; - the follow-up tool name; - the complete JSON payload is written under the active run directory: `tool_results/*.json`; - agents use `mcp__evaluation-tools__read_tool_result` to read bounded chunks of a stored result by `offset` and `max_chars`. This gives Praxist character-window pagination without requiring a separate cursor protocol for every tool: ```text summary tool response -> _tool_output.full_result_ref -> read_tool_result(ref, offset, max_chars) ``` The mechanism is lossless for system-generated tool data: inline summaries may be truncated, but the full JSON remains in the run artifacts. Current scope: - `evaluation-tools.get_leaderboard`; - `frontier-tools.get_frontier`; - `finding-graph-query.get_finding_neighbors`; - `finding-graph-query.get_finding_subgraph`; - `finding-graph-query.get_unlinked_recent_findings`; - `evaluation-tools.read_tool_result`. Out of scope: - built-in agent CLI `Bash`, `Read`, and shell command output. Praxist cannot hard-cap tools built into the agent runtime from a generic MCP tool server. Prompts should still instruct agents to avoid `cat` on large JSON/log files and to prefer compact Praxist tool responses. ## Task-Local Evaluation Toolification Problem: expensive tasks can give peers long prompt instructions for manually launching staged benchmarks, waiting for files, parsing benchmark JSON, applying gates, and deciding whether to escalate. Agents often copy commands, inspect large raw JSON files, or repeat shell checks. Design: - keep task-specific evaluation mechanics inside the external task project; - provide one compact public evaluator entrypoint under the task's `evaluations/` directory; - keep Praxist core and generic plugins unchanged; - have the peer implement a variant, then run one task-local command; - preserve raw benchmark JSON and logs under the run results directory; - print only a compact gate summary to stdout. Example command shape: ```bash python "$PRAXIST_TASK_PROJECT_PATH/evaluations//run.py" \ --variant-path "" \ --output-dir "/" \ --data-dir "" \ --max-stage "" ``` The task-local tool owns benchmark invocation, promotion gates, raw evidence preservation, concise stdout summaries, and exact maturity telemetry such as `effort_ratio` and `coverage_ratio`. For smoke runs, expose a task-owned argument or environment variable that limits the evaluator stage without changing Praxist core behavior. ## Measure The Effect Savings depend on runtime caching, provider metering, evaluator output size, event cadence, and agent behavior. Praxist does not promise a fixed token or billing reduction. Compare equivalent runs using canonical usage artifacts. Record input, cached input, uncached input, output, session count, sessions per peer-generation, and cache-hit ratio. Also verify task correctness and artifact recoverability; a lower token count is not useful if peers repeat work or miss evidence. Use `praxist-diagnostic` to report input, cached input, uncached input, output, session count, sessions per peer-generation, and cache-hit ratio. A high cache hit rate can coexist with excessive logical token use when every new thread repeats broad bootstrap reads. ## Testing Expectations Changes in this area should include: - unit tests for result-ref storage, path confinement, chunk reads, and list truncation metadata; - tool adapter tests proving original business fields still exist; - task-local tests for evaluator gates and reuse-existing summaries; - offline integration tests proving startup exposes the declared tool names; - API provider gating tests proving direct DeepSeek behavior is unchanged; - event tests proving finding bursts coalesce while lifecycle/resource events remain immediate; - relay tests proving only OpenRouter receives sticky-session metadata; - docs build after guide or docstring changes. Do not test exact prompt prose or temporary output ordering. Test stable contracts: bounded inline output, full-result recoverability, and task-local gate decisions. --- # Workflow Stages Workflow stages are executable steps in a Praxist run. ## Research Loop `workflow_stage:research_loop` is mandatory. It owns the peer cohort, shared findings, frontier, finding-graph guidance, Principal Investigator (PI) and Chair synthesis, prompt layout, generation boundaries, and run artifacts. The stage also owns the Deep Innovation Gate (DIG), a generation-scoped pre-code design process, and the independently configurable Quality-Diversity (QD) allocation path. See [Deep Innovation Gate](https://praxist.sapient.inc/en/docs/guides/deep-innovation-gate) and [Quality-Diversity Allocation](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator). Each generation materializes its executable topology in `gen_/research_topology.json` before running the standard parallel peer cohort. See [Research Topology Audit API](https://praxist.sapient.inc/en/docs/guides/research-topology-and-module-api). ## Interface Placeholders `workflow_stage:ideation_stub` and `workflow_stage:paper_writing_stub` are registered interface placeholders, not product modules. They remain disabled by default and do not provide ideation or paper-writing workflows. ## Local Reviewer `workflow_stage:reviewer_stub` provides an optional local artifact and provenance review when explicitly run in one of these modes: ```text local artifact artifacts run_artifact claim_check review ``` The reviewer reads `artifact_index.jsonl`, `trajectory.jsonl`, and `run_summary.json`, verifies artifact hashes and references, and writes `workflow/reviewer_report.json`. It does not rerun evaluators, assess scientific quality, or affect frontier, incubator, Gems, or leaderboard state. It refuses to append after `run.finalized`, preserving that event as the trajectory terminus. ## Stage Contract An executable stage must: - validate its input contract; - request budget before expensive work; - emit lifecycle events; - write replayable artifacts; - preserve partial outputs where safe; - report terminal status. Stage semantics belong in Python workflow plugins, not shell wrappers or task harnesses. --- # Tool Servers Tool servers are generic plugins that expose bounded capabilities to peers and panels. Task projects decide when a capability is scientifically relevant; task-specific instructions do not belong in the server. ## Catalog | Tool server | Purpose | |---|---| | `evaluation_tools` | Compact evaluator and leaderboard access | | `frontier_tools` | Committed frontier views | | `memory_tools` | Peer-memory lookup | | `finding_graph_query` | Advisory finding-graph queries | | `prior_work_tools` | Existing-work lookup | | `run_report` | Human-readable derived reports | | `literature_lookup` | Public scientific literature/database context | | `existing_mcp_tools_shim` | Compatibility bridge for declared external tools | The finding graph is built by a graph-maintainer plugin; its tool server is only the query surface. ## Specialized Contracts [Scientific Literature and Database Lookup](https://praxist.sapient.inc/en/docs/guides/scientific-literature-lookup) owns source coverage, provenance, current-environment limits, and runtime behavior for `literature_lookup`. [Run Reports](https://praxist.sapient.inc/en/docs/guides/user-facing-reports-and-init) owns automatic/manual report triggers, structure, metric interpretation, and canonical-state boundaries for `run_report`. ## Frontier Views `frontier_tools.get_frontier` reads committed membership from `frontier/frontier_manifest.json`; it does not rerun promotion. The effective task-spec maturity policy remains available for interpretation, but a live task file cannot replace the run's frozen policy. The latest-generation view uses current compact lanes. Historical cutoffs are reconstructed from the canonical per-generation ledger so later capacity eviction does not rewrite earlier membership. Compact state without immutable history, such as unverifiable historical Gems membership, is omitted and marked incomplete rather than guessed. The response reports canonical and returned counts, categorized skips, policy source, and `frontier_view_integrity_status`. If view construction hides every entry from a non-empty canonical lane, it returns `canonical_entries_hidden` as an integrity error. Validation candidates remain separate and are never promoted by the reader. ## Tool Conformance Tool-server tests cover manifest resolution, handler construction, allowed tool names, bounded normalized output, failure normalization, redaction, and missing optional credentials. Full payload recovery and inline output limits are defined in [Cost Optimization](https://praxist.sapient.inc/en/docs/guides/cost-optimization#tool-output-limits-and-full-lookup). --- # Budget Policies Budget is dynamic in Praxist. It is not a fixed tuple copied once into a run and then blindly enforced for every experiment. ## Budget Requests Agents and workflow stages may request budget in the units accepted by the current core validator: - `tokens`; - `wall_clock_seconds`; - `gpu_hours`. Requests should include scope, reason, estimated cost, expected value, and the action that will consume the budget. ## Policy Decisions A BudgetPolicy may: - auto-grant low-risk requests; - downscope a request; - ask a Principal Investigator (PI) or Chair planning agent to review unusually large requests; - deny requests that would damage the run or exceed operator limits. The default posture is result preservation. A promising experiment should be allowed to finish when it is inside a reasonable envelope and can produce useful artifacts, even if exact metering is imperfect. ## Usage Records Usage records can be exact, estimated, partial, or unknown. Unknown usage must be recorded explicitly as `usage_unknown` instead of `0`. Late accounting failure should warn and preserve findings when possible. ## Budget Tests Budget policy tests should cover grant, deny, downscope, review routing, usage_unknown, replay visibility, and behavior when a peer produces results before usage accounting completes. --- # Agent OOBE Runbook This runbook is the machine-facing contract for a Codex- or Claude Code-managed Praxist installation. It keeps the agent-managed lane separate from the local terminal wizard while reusing the same setup profiles, configuration files, and validation. ## Boundary Use this runbook when the operator asks Codex or Claude Code to install and configure Praxist from PyPI, or explicitly points the agent at a source checkout for development. Do not use it for an ordinary runtime repair after OOBE. An installation request does not authorize project selection or research launch. Stop after setup and readiness; takeover is a separate action after the operator reads [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task). Set one conversation-local skill host, `codex` or `claude`, from the interface running this workflow. [Agent Skills](https://praxist.sapient.inc/en/docs/user-guide/skills#install-or-refresh) owns the installation locations and refresh commands. The skill host is only the human interaction surface. It does not choose or change the Praxist peer runtime selected by the setup profile. Keep these choices in the current conversation until setup completes: - selected Praxist executable; - selected setup profile ID. Do not create an OOBE state file. Legal-terms acceptance, installed configuration (including the explicitly selected profile ID), product-usage consent, `praxist doctor`, and task artifacts are the durable sources from which interrupted setup is resumed. ## Interaction Contract 1. Prefer the current agent's native structured choice UI. If it is unavailable, ask the operator to run the matching local TTY selector. Do not ask for typed `yes`/`no` answers. 2. Use Up/Down and Enter for local choices. Treat Esc as back/cancel without undoing a completed package installation. 3. Never ask the operator to paste an API key into chat. API providers must use the local masked input opened by `praxist setup`; it displays one `*` per character and keeps the key out of chat, argv, logs, and shell history. 4. Ask for a path only when bounded discovery cannot identify the intended project. Never scan the whole home directory, a storage mount, or a dataset tree. 5. Do not use `--force-unmanaged` automatically. If bundled skill names collide with operator-owned paths, show the complete conflict list and offer: keep the existing skills, back them up and replace them, or cancel. Continue only with the operator's selected action. 6. Never accept the Praxist Fair Source License or User Agreement on the operator's behalf or infer acceptance from an installation request. Keep legal acceptance separate from optional product-usage consent. 7. A usable API provider default, saved login, exported key, or successful doctor result is not a profile choice. Only the operator may choose the profile. ## Workflow 1. **Preflight and install.** Verify Python 3.11+ and respect the operator's active or explicitly selected Python environment. Do not replace a task environment or modify system Python. Install the package and maintained runtime integrations without selecting an API provider on the operator's behalf: ```bash python3 -m pip install --index-url https://pypi.org/simple "praxist[agents,codex]" ``` If no suitable writable environment is active, create a dedicated virtual environment first and keep using its Python and `praxist` entrypoint for the rest of OOBE. Do not silently install into an externally managed interpreter. A successful pip command proves only that the distribution is installed. Legal terms, privacy, runtime, skills, writable examples, and readiness remain pending. Continue with the steps below; do not report OOBE completion at this boundary. Never launch a read-only source/package-resource example copy. Immediately run `praxist setup --agent-managed`. This read-only command is the machine-owned decision checkpoint. Follow its `next_required_action` and rerun it after each completed decision. 2. **Legal terms.** Run `praxist user-agreement status --json`. If the current legal bundle is not accepted, present three native choices with no acceptance preselected: review the complete terms, agree and continue, or cancel setup. For review, surface the canonical root `LICENSE.md`, the two packaged Markdown files under `docs/legal/`, and the `license_url` and `review_url` returned by the status command as scrollable links. Use `praxist user-agreement review --print` only when neither interface is available, because printing the full text into chat is the least usable fallback. Repeat the choices after review. Explain that the Fair Source License includes eligibility, revenue-threshold, attribution, distribution, and use restrictions. Only after the operator explicitly agrees to the complete bundle may the Agent run: ```bash praxist user-agreement accept --agent-reply Agree ``` Verify that status now reports `accepted: true`. A cancellation ends OOBE without changing runtime configuration. 3. **Privacy.** Run `praxist product-usage status --json`. Legal acceptance is not optional telemetry consent. If collection is available and consent is unset, offer review, share, and skip choices with neither sharing nor skipping preselected. Review uses the packaged `docs/legal/product-usage-data-notice.md` or its hosted page rather than dumping the notice into chat. Record only the explicit selection with `praxist product-usage consent --agent-reply `. If no choice is made, leave consent unset. When `collection_available` is false, explicitly tell the operator that this build collects no product-usage data and no privacy authorization is required. 4. **Runtime.** Read the machine-owned options from `praxist setup --agent-managed` (or `--list-profiles`). Present those complete setup profiles as selectable API provider, agent runtime, concrete model, and authentication combinations. Apply a profile only after the operator chooses it. Before applying it, state the profile's `authorization_detail`: Codex-native uses the saved ChatGPT/Codex login without an API provider key or new authorization code; API-backed setup profiles require the matching key through local masked input. Then run `praxist setup --profile --install-skills `, where `` is `codex` or `claude` from the current interface. Run the command in a local interactive terminal when the profile needs either an API key or a Codex-native ChatGPT login. The latter uses and verifies the SDK-pinned Codex binary; no other profile may require that login. If secure local interaction is unavailable, print the exact setup command for the operator and wait; never route credentials through the agent conversation. `praxist setup --interactive` handles any same-name skill conflict locally with keep, backup-and-replace, and cancel choices. Non-interactive setup refuses to overwrite an operator-owned path. Rerun `praxist setup --agent-managed` and require `profile.selected: true`; a configured API provider default without that confirmation remains incomplete. 5. **Readiness and stop.** Require `setup_decisions_complete: true`, run the matching host diagnostics, and report the selected profile, installed skill host, writable example locations, and any concrete blocker. The final `next_required_action` is `run_doctor_then_finish_setup`. Do not discover or select a research project, invoke a takeover skill, or launch a run. Link the operator to [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task) for the later, separate takeover workflow. Finish when installation, explicit setup decisions, skill registration, and readiness checks are complete, or when one concrete blocker remains. --- # Contributing This repository is edited by both humans and agents. The durable rule is simple: keep system code, generic plugins, task templates, complete examples, and real task projects physically separate. ## Read First - `AGENTS.md` is the machine-friendly repository contract. - `docs/concepts/architecture.md` is the active architecture overview. - `docs/index.md` is the source documentation entry. ## Default Change Flow 1. Identify whether the change belongs to core, a generic plugin, a task template, a complete example, an external task project, tests, docs, or operator scripts. 2. Make the smallest change that fits the existing boundary. 3. Add or update tests at the same boundary. 4. Rebuild docs when docstrings or docs changed. 5. Record high-risk architectural implementation context in the commit or pull-request description. ## Stable Docstrings Public and semi-public Python APIs use Google-style docstrings. A docstring should describe the contract that future callers can rely on, not the history of one bug fix. Use comments for non-obvious invariants, recovery behavior, and failure policy. Do not add comments that only restate the next line of code. ## Default Verification ```bash uv sync --group dev --extra docs uv run python -m unittest discover -s tests -q uv run python scripts/run_test_coverage.py unit --fail-under 90 --fail-under-statements 95 uv run python scripts/run_test_coverage.py integration uv run python -m compileall -q praxist tests templates examples scripts uv run python scripts/build_docs_site.py git diff --check ``` Narrow changes can start with narrow tests, but the default handoff should still include the full suite above. The coverage command writes ignored local reports under `cover/unit/` and `cover/integration/`. The `unit` profile is the offline non-integration test layers and is held at 90% branch-aware total coverage plus 95% statement coverage; the `integration` profile is `tests/integration` and remains observational. It uses `coverage.py` from the dev dependency group and keeps the test runner on `unittest`. ## What Requires Extra Care Extra care is required when changing startup, plugin resolution, task path resolution, credential selection, runtime invocation, prompt layout, event-driven peer scheduling, finding graph guidance, budget policy, replay verification, or run artifact schemas. Those changes should include focused tests and, when they change architecture contracts, a concise rationale in the pull-request description. --- # Configuration Discipline > Core and plugin domain code consume explicit configuration objects. Ambient > environment variables are read at operator entry boundaries, then resolved > once into frozen runtime configuration. `RunConfig` centralizes run-critical settings. Narrow compatibility paths still read ambient variables, and selected launcher or scheduler boundaries may pass documented host values into child processes. This contract keeps startup, replay, API provider routing, credential handling, and agent runtime execution deterministic. ## Configuration Flow ```text CLI entry argparse + selected environment values | v frozen RunConfig / domain contracts | v core + resolved plugins | v runtime-owned SDK/child-process environment ``` Configuration priority is: ```text explicit CLI > explicit environment > override spec > task defaults ``` A value is normalized once at startup. Downstream code receives `RunConfig` or a narrower dataclass such as `AgentRunRequest`, `ModelCallSpec`, `CredentialRef`, or `RuntimeSandboxIntent`. It must not reconstruct API provider, model, task, agent runtime, or credential choices from strings later in the call graph. ## Ingress Boundaries CLI entrypoints under `praxist/cli/` and `praxist/run.py` may read documented Praxist configuration variables. They are responsible for: - parsing operator intent; - applying precedence; - resolving task and run paths; - selecting agent runtime/API provider/model refs; - passing raw credentials only to credential resolution; - producing frozen configuration for downstream consumers. Task projects do not own research startup. The Praxist CLI resolves the task, run directory, provider, credentials, plugins, budget, and workflow before launch; task-specific runtime values enter through the validated task contract. ## Configuration Boundary Code under `praxist/core/`, generic plugin domain logic, and infrastructure services must receive configuration explicitly. Direct ambient reads are not a transport mechanism between layers. The following patterns are out of contract: - reading `PRAXIST_*` values in core business logic; - import-time constants populated from environment variables; - mutating `os.environ` so another in-process layer can discover a value; - inferring an API provider from model-name punctuation after startup resolved it; - copying all host environment variables into a runtime or tool process; - passing raw credentials through task files, prompts, logs, or trajectory. ## Runtime Egress A runtime adapter may construct an SDK/client or child-process environment from its explicit execution context. This is egress, not a second configuration resolution pass. The adapter must: - include only the selected API provider credential; - redact credentials from events, errors, and logs; - pass non-secret task runtime values needed by shell or MCP children; - avoid forwarding unrelated host secrets; - keep runtime-private state under the selected run directory; - preserve timeout, cancellation, and sandbox intent from the request. For `agent_runtime:codex_sdk`, the official SDK owns the local app-server. The runtime creates a private run-scoped relay only for supported Chat Completions API providers and attaches selected Praxist tools directly over MCP. The human-facing Codex CLI is a separate operator surface. When native OpenAI is selected and no API key exists, its saved ChatGPT authentication may authorize the SDK runtime; personal plugins, skills, hooks, MCP servers, instructions, and runtime state do not become peer configuration. ## Credentials Credential resolution converts a raw secret into a redacted `CredentialRef` plus the minimum runtime material required for the selected API provider. Core does not inspect API provider keys. An agent runtime may inject the resolved key into an SDK or private child process when that external interface requires an environment variable. A runtime may also expose an optional managed-credential discovery hook. Core accepts only a redacted `CredentialRef` whose `scope` and `provider` match and whose `target_ref` is absent or matches the selected API provider ref. Environment credentials win, and resolve-only startup never performs an agent runtime authentication probe. API provider failover, cooldown, and key selection remain Python control-plane semantics. Shell wrappers and task projects must not duplicate them. ## Replay And Audit Persisted startup configuration records the resolved non-secret values used by the run. Replay and diagnostics read that frozen state rather than trying to reconstruct the parent shell environment. This provides: - deterministic comparison between runs; - one explanation for API provider/model/agent runtime selection; - stable cost and usage attribution; - a bounded secret-review surface; - reproducible runtime and task-environment setup. Unknown or unavailable values must remain explicit (`usage_unknown`, missing optional capability, or a warning). Do not convert unknown state to a false zero or infer it from unrelated environment variables. ## Task Experiment Configuration Task evaluators own the scientific treatment they execute. When launch-time arguments, environment overrides, protocol settings, or task-local config files can change a result without changing variant code, the existing result summary may publish a secret-free top-level `effective_config` object and an explicit `effective_config_complete` boolean. Praxist hashes that object and propagates only its digest, status, and the existing result-summary path; the summary remains the single owner of the full configuration. The evaluator must publish resolved treatment values after applying defaults and parsing, not a raw environment snapshot. An omitted setting and an explicit setting equal to its resolved default are the same configuration. An evaluator claiming an exact replication should also publish the parent's `replication_of_effective_config_sha256`. Praxist reports a match only when the current configuration is complete and its digest equals the parent digest. Missing provenance does not change ordinary evidence maturity, promotion, closing, or legacy evaluator behavior; it only means the run cannot support an exact-replication claim. ## Verification Tests should inject configuration through constructors, CLI arguments, or a bounded environment mapping at the entry boundary. They should also verify that: - serialized configuration is redacted; - runtime child environments exclude unrelated secrets; - API provider/model normalization happens once; - task runtime values reach MCP/shell children through explicit context; - replay does not depend on the ambient environment; - exact-replication fixtures distinguish identical code under different effective task configurations; - new core/plugin code does not introduce undocumented environment reads. --- # Runtime Model This page explains the research loop at the agent-session boundary. It does not define API provider routing, resource scheduling, evidence policy, or budgets; those contracts have dedicated guides. ## Agent Sessions A peer opens an agent session with a rendered prompt layout and a normalized runtime request. The request identifies the model profile, credential reference, tools, sandbox intent, cache policy, and timeout. The selected runtime turns SDK events into the common Praxist event/result protocol. After a session returns, the peer waits for a meaningful event instead of immediately resending the same context. Examples include a revised shared finding, lifecycle signal, resource-supply event, or heartbeat expiry. Event cadence and continuation behavior are runtime policies; canonical findings and task state remain on disk. [Agent Runtimes](https://praxist.sapient.inc/en/docs/guides/agent-runtimes) owns adapter behavior and [Cost Optimization](https://praxist.sapient.inc/en/docs/guides/cost-optimization) owns lossless finding-event batching. ## Prompt Layout PromptLayout V1 separates: - **frozen blocks** shared across compatible calls; - **semi-static blocks** changed by task or role contracts; and - **dynamic blocks** containing generation, frontier, memory, and run state. The rendered prompt and its layout manifest are audit snapshots. They preserve what the agent saw for replay, but later stages regenerate views from current canonical result, finding, frontier, Gems, memory, and boundary state. Peer memory is a navigation index into those artifacts, not a replacement for them. [Peer Memory](https://praxist.sapient.inc/en/docs/guides/peer-local-structured-memory-long-context) defines its lifecycle. ## Related Contracts - [Central Experiment Scheduler](https://praxist.sapient.inc/en/docs/guides/central-resource-scheduler): experiment admission, process ownership, and resource release. - [Budget Policies](https://praxist.sapient.inc/en/docs/guides/budget-policies): budget decisions and usage records. - [Cost Estimation](https://praxist.sapient.inc/en/docs/guides/costs): interpretation of measured usage. - [Research Loop](https://praxist.sapient.inc/en/docs/guides/research-loop-variant-generation-flow): full generation sequence and artifact inheritance. --- # Generic Plugins Generic plugins are reusable system components under `praxist/plugins/**` or another explicit plugin root. A task project may also ship its own generic plugins under `/.praxist/plugins/`; those are discovered as `source="task_project"` only when `--task-path` selects the task (see [Task Projects → Boundary Rules](https://praxist.sapient.inc/en/docs/guides/task-projects#boundary-rules)). ## Plugin Manifest Each plugin has a `plugin.yaml` manifest describing: - plugin kind and name; - version and stability; - executable entrypoint or manifest-only contract; - dependencies and compatibility; - declared code/assets for replay hash coverage. The plugin loader discovers candidates, resolves dependencies, checks source priority, and writes the selected plugin manifest into the run directory. ### Stability as an interface contract `stability` describes the **interface contract** a plugin promises — schema, prompt shape, role-API backward compatibility — not whether the plugin is trustworthy to execute. Trust comes from the plugin's **source**, gated via `TRUSTED_EXECUTION_SOURCES` (`bundled` and `task_project`). The two paths are gated differently: - **Bundled plugins** must declare the kind's strict expected stability (e.g. `v1_stable` for `panel_topology`, `agent_runtime`, `workflow_stage`). A bundled plugin affects every task project, so the contract there has to hold. - **Task-local plugins** under `/.praxist/plugins/` are scope-isolated to one task project and high-churn by design; they may declare any `stability` value (commonly `v0_experimental`) without triggering the kind-mismatch check. Source trust is enough. ## Minimal Executable Plugin An executable plugin usually has this shape: ```text praxist/plugins/// plugin.yaml adapter.py README.md # optional, for complex plugins ``` `plugin.yaml` should declare the plugin ref, compatibility, entrypoint, and code files that participate in source hashing. `adapter.py` should expose a small factory or adapter object matching the kind-specific contract. The plugin content hash covers the manifest and its declared code and assets. Imported modules or assets omitted from the manifest are outside that plugin content hash. Do not name every implementation file `plugin.py` by habit. Use names that describe the plugin's internal structure. ## Plugin-Supplied Assets Some plugin kinds accept assets shipped alongside the manifest, declared through dedicated manifest fields rather than ad-hoc paths. - `panel_topology`: a plugin may declare `topology.prompts_dir` to ship its own Jinja prompt templates for `BasePI` and `ChairArbiter`. The bundled prompts for the multi-agent Principal Investigator (PI) panel are used as a fallback for anything the plugin does not override. See [Panel Topology Prompts](https://praxist.sapient.inc/en/docs/concepts/panel_topology_prompts). When a plugin ships assets, declare the paths under `code` / `assets` in `plugin.yaml` so they participate in replay source hashing. ## What Belongs in a Plugin Put code in a generic plugin when it is reusable across task projects: - runtime adapters; - API provider adapters; - tool servers; - workflow stages; - graph maintainers; - generic budget policies. Do not put benchmark-specific research facts or task-local role contracts into a generic plugin. Those belong in the task project. ## Decision Test Before adding a plugin, ask whether two unrelated task projects could use it without copying task facts. If the answer is no, it probably belongs in a task project. Before adding code to core, ask whether the behavior can be selected, replaced, or disabled through a plugin. If yes, it belongs in a plugin. ## Plugin Tests Plugin changes should add: - plugin-local unit tests when the plugin has real code; - kind-specific conformance tests under `tests/conformance/`; - workflow smoke tests when the plugin participates in startup or run execution; - replay/hash tests when plugin code or assets affect run reproducibility. The default test path must not require real model keys, network, GPUs, or an external task checkout. --- # Research Topology Audit API Praxist records the executable research topology for each generation and exposes a structured, read-mostly API for external modules. These surfaces are audit and integration boundaries; they do not replace findings, frontier, Gems, memory, or result artifacts. ## Executed Topology The bundled executor runs one parallel cohort per generation: ```text generation -> peer cohort -> findings -> generation boundary -> frontier / Gems / memory / PI and Chair planning ``` Before the cohort starts, Praxist writes: ```text gen_/research_topology.json ``` The sidecar contains a `ResearchTopologySpec` with worker nodes, edges, policy, and metadata. It records what the executor is about to run. Generic worker types in the schema are descriptive vocabulary; the bundled executor does not claim to execute a worker type unless that node appears in the materialized topology. ## Module API `ResearchLoopModuleAPI` provides these operations: ```text submit_recommendation request_topology_change list_commands list_findings get_frontier_summary get_validation_signals get_gems_summary get_memory_summary get_run_status ``` `get_frontier_summary` returns durable frontier evidence. `get_validation_signals` returns compact task-defined signals for triage and follow-up planning. Validation signals retain their actual stage and coverage; they do not become clean parents unless the task contract explicitly grants that authority. ## Command Queue Recommendations and topology-change requests are appended to: ```text external_requests/research_commands.jsonl ``` The bundled executor records these commands but does not inject them into peer prompts or mutate a live topology. A queued command is operator intent, not an experiment contract. PI and Chair planning continue to own executable peer assignments. ```text external module -> command queue -> audit and operator review ``` External modules should use this API instead of editing prompts, agendas, or generation artifacts directly. ## Boundary Rules - `GenerationLoop` owns lifecycle, resume, and generation boundaries. - The topology executor owns cohort execution behind that boundary. - Topology policy and worker adapters belong in plugins, not task-specific core branches. - Requests use the current `queue_for_generation_boundary` policy. - Queued requests have no scientific or promotion authority by themselves. --- # CLI Reference This page is generated from the same `argparse` tree used by the `praxist` executable. The command implementation is the sole definition; rebuild the documentation after changing CLI arguments. ```text usage: praxist [-h] [--version] ... ``` ## Global arguments | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | | `--version` | no | show program's version number and exit | ## Commands | Command | Purpose | |---|---| | [`praxist configure-llm`](#praxist-configure-llm) | Persist a built-in Praxist LLM provider profile. | | [`praxist docs`](#praxist-docs) | Open or print the hosted Praxist documentation. | | [`praxist doctor`](#praxist-doctor) | Check Praxist host readiness. | | [`praxist examples`](#praxist-examples) | List or install complete writable example projects. | | [`praxist install-skills`](#praxist-install-skills) | Install bundled Praxist skills for Codex or Claude Code. | | [`praxist uninstall-skills`](#praxist-uninstall-skills) | Remove Praxist-managed agent skill registrations. | | [`praxist product-usage`](#praxist-product-usage) | Review or change pseudonymous product-usage consent. | | [`praxist monitor`](#praxist-monitor) | Watch Praxist run state in a live read-only terminal dashboard. | | [`praxist resolve`](#praxist-resolve) | Resolve a task project's plugin manifest (no LLM calls). | | [`praxist resume`](#praxist-resume) | Resume an interrupted Praxist run. | | [`praxist setup`](#praxist-setup) | Configure this host for Praxist operation. | | [`praxist start`](#praxist-start) | Launch a new Praxist research run (registry-backed). | | [`praxist status`](#praxist-status) | List known Praxist experiment runs. | | [`praxist stop`](#praxist-stop) | Stop a Praxist run by run_id, or stop everything with --all. | | [`praxist takeover`](#praxist-takeover) | Open Codex or Claude Code and hand off a project to Praxist takeover. | | [`praxist uninstall`](#praxist-uninstall) | Remove the user-level Praxist installation. | | [`praxist user-agreement`](#praxist-user-agreement) | Review the Praxist License and User Agreement or inspect acceptance status. | ## `praxist configure-llm` Persist a built-in Praxist LLM provider profile. ```text usage: praxist configure-llm [-h] --provider PROVIDER [--model MODEL] [--agent-system {claude_sdk,codex_sdk}] [--api-key-stdin | --api-key-env API_KEY_ENV | --no-api-key | --remove-api-key] [--config-file CONFIG_FILE] [--project-env-file PROJECT_ENV_FILE] [--print-source-command] [--json] [--dry-run] ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | | `--provider` | yes | Built-in provider name or compatible provider plugin reference. | | `--model` | no | Provider model name to persist. | | `--agent-system` | no | Agent runtime selection to persist. Choices: `claude_sdk`, `codex_sdk`. | | `--api-key-stdin` | no | Read the provider API key from stdin; a local TTY shows one * per character. | | `--api-key-env` | no | Read the provider API key from this environment variable. | | `--no-api-key` | no | Update non-secret provider settings without writing an API key. | | `--remove-api-key` | no | Remove this provider's stored API key from the selected config file(s). | | `--config-file` | no | Config file to update (default: $PRAXIST_CONFIG_FILE or the user config). | | `--project-env-file` | no | Also write Praxist LLM config to this explicit task-local .env file. | | `--print-source-command` | no | Print the shell command that loads the selected config file. | | `--json` | no | Emit the result as JSON. | | `--dry-run` | no | Validate and report changes without writing files. | ## `praxist docs` Open or print the hosted Praxist documentation. ```text usage: praxist docs [-h] [--no-open] ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | | `--no-open` | no | Print the documentation URL without opening a browser. | ## `praxist doctor` Check Praxist host readiness. ```text usage: praxist doctor [-h] [--json] [--task-path TASK_PATH] [--config-file CONFIG_FILE] [--agent-system {claude_sdk,codex_sdk}] [--model-provider MODEL_PROVIDER] [--model MODEL] [--codex-native] [--target {auto,codex,claude}] [--advisory] ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | | `--json` | no | Emit the readiness report as JSON. | | `--task-path` | no | Also validate this task project and its runtime environment. | | `--config-file` | no | Config file to inspect (default: $PRAXIST_CONFIG_FILE or the user config). | | `--agent-system` | no | Check one research runtime (default: configured runtime or claude_sdk). Choices: `claude_sdk`, `codex_sdk`. | | `--model-provider` | no | Check one model_provider ref using the same precedence as praxist start. | | `--model` | no | Check this selected model (Codex-native verifies it in the account catalog). | | `--codex-native` | no | Check codex_sdk with native OpenAI and the saved ChatGPT login. | | `--target` | no | Check bundled skills for this agent host (default: detect managed installs). Choices: `auto`, `codex`, `claude`. Default: `auto`. | | `--advisory` | no | Always return exit 0 while retaining readiness failures in the report. | ## `praxist examples` List or install complete writable example projects. ```text usage: praxist examples [-h] ... ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | ## `praxist install-skills` Install bundled Praxist skills for Codex or Claude Code. ```text usage: praxist install-skills [-h] [--target {codex,claude}] [--target-dir TARGET_DIR] [--mode {copy,symlink}] [--replace] [--force-unmanaged] [--migrate-legacy-symlinks] [--dry-run] [--json] ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | | `--target` | no | Skill host. Default: codex. Choices: `codex`, `claude`. Default: `codex`. | | `--target-dir` | no | Override the target skill directory. | | `--mode` | no | Register skills by copying package content or linking a source checkout. Choices: `copy`, `symlink`. Default: `copy`. | | `--replace` | no | Refresh existing Praxist-managed entries; unmanaged paths require --force-unmanaged. | | `--force-unmanaged` | no | With --replace, back up and replace unmanaged entries whose names exactly match bundled Praxist skills. Unrelated skills are untouched. | | `--migrate-legacy-symlinks` | no | With --replace, explicitly adopt old Praxist repo-style symlinks that predate the ownership manifest. | | `--dry-run` | no | Report actions without changing the target directory. | | `--json` | no | Emit the result as JSON. | ## `praxist uninstall-skills` Remove Praxist-managed agent skill registrations. ```text usage: praxist uninstall-skills [-h] [--target {codex,claude}] [--target-dir TARGET_DIR] [--dry-run] [--json] ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | | `--target` | no | Skill host. Default: codex. Choices: `codex`, `claude`. Default: `codex`. | | `--target-dir` | no | Override the target skill directory. | | `--dry-run` | no | Report removals without changing the target directory. | | `--json` | no | Emit the result as JSON. | ## `praxist product-usage` Review or change pseudonymous V2 product-usage consent. Withdrawal stops future capture and deletes unsent local events; delivered events expire through scheduled retention. ```text usage: praxist product-usage [-h] ... ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | ## `praxist monitor` Render a live read-only dashboard from praxist status, orchestrator snapshots, peer memory health, recent logs, and lightweight host load. The dashboard runs directly in the current terminal and never controls the Praxist research process. ```text usage: praxist --monitor [-h] [--run-id RUN_ID] [--run-dir RUN_DIR] [--task-path TASK_PATH] [--latest] [--interval INTERVAL] [--once] [--follow] [--no-clear] [--plain] [--log-lines LOG_LINES] [--peer-limit PEER_LIMIT] ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | | `--run-id` | no | Monitor one run id. | | `--run-dir` | no | Monitor one run dir. | | `--task-path` | no | Prefer active rows for this task path. | | `--latest` | no | Select the latest active run row when more than one exists. | | `--interval` | no | Frame interval in seconds (default: 0.2 for the fullscreen TUI, 1 for plain text). | | `--once` | no | Render one frame and exit. | | `--follow` | no | Keep refreshing even when stdout is not an interactive terminal. | | `--no-clear` | no | Append frames instead of clearing the terminal between refreshes. | | `--plain` | no | Use the legacy plain-text monitor instead of the fullscreen TUI. | | `--log-lines` | no | Recent log lines to show for the selected run (default: 18). Default: `18`. | | `--peer-limit` | no | Maximum peer rows to render (default: 24). Default: `24`. | ## `praxist resolve` Discover and resolve a task project's plugin manifest without making any LLM calls. Equivalent to: python -m praxist.run run --task-path --resolve-only --local Exits non-zero on resolution failure (manifest schema error, missing plugin, etc.). On success, emits a JSON document on stdout summarizing the resolved run identity. ```text usage: praxist resolve [-h] [--config-file CONFIG_FILE] [--agent-system {claude_sdk,codex_sdk}] [--workspace WORKSPACE] [--run-dir RUN_DIR] [--runtime RUNTIME] [--codex-native] [--model-provider MODEL_PROVIDER] [--budget-policy BUDGET_POLICY] [--credential-profile CREDENTIAL_PROFILE] [--model MODEL] [--result-summary RESULT_SUMMARY] [task_path] ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | | `task_path` | no | Path to the task project directory (default: current directory). Default: `.`. | | `--config-file` | no | Config file to load (default: $PRAXIST_CONFIG_FILE or the user config). | | `--agent-system` | no | Agent system used to resolve runtime/provider defaults. Choices: `claude_sdk`, `codex_sdk`. | | `--workspace` | no | Workspace directory (default: current working directory). Default: ``. | | `--run-dir` | no | Override run artifact directory. Defaults to the task project's runtime_outputs.root / experiments directory; paths inside the Praxist source checkout are rejected. Default: ``. | | `--runtime` | no | Override agent_runtime plugin ref. Default: ``. | | `--codex-native` | no | Resolve with codex_sdk, native OpenAI, and saved ChatGPT login while ignoring API-key/custom-endpoint configuration. | | `--model-provider` | no | Override model_provider plugin ref. Default: ``. | | `--budget-policy` | no | Override budget_policy plugin ref. Default: ``. | | `--credential-profile` | no | Override credential profile name (rarely needed for resolve-only). Default: ``. | | `--model` | no | Override agent model. Not used by resolve-only itself, but propagated for parity. Default: ``. | | `--result-summary` | no | Validate one evaluator-produced JSON summary against the task's maturity telemetry contract before resolving. Default: ``. | ## `praxist resume` Continue an existing Praxist run directory from its last safe completed generation boundary. The target may be a registry run_id from ``praxist status`` or a direct experiments/run_* path. ```text usage: praxist resume [-h] [--task-path TASK_PATH] [--config-file CONFIG_FILE] [--agent-system {claude_sdk,codex_sdk}] [--runtime RUNTIME_REF] [--codex-native] [--model MODEL] [--model-provider MODEL_PROVIDER_REF] [--strategy {auto,mixed,explore,exploit}] [--cohort COHORT] [--generations GENERATIONS] [--server] [--daemonize] [--resume-policy {completed_generation}] [--force] [--startup-timeout STARTUP_TIMEOUT] [--json] target ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | | `target` | yes | Registry run_id or path to an existing Praxist run directory. | | `--task-path` | no | Override task project path when resuming from a run directory. | | `--config-file` | no | Config file to load (default: $PRAXIST_CONFIG_FILE or the user config). | | `--agent-system` | no | Override agent system for the resumed launch. Choices: `claude_sdk`, `codex_sdk`. | | `--runtime` | no | Override agent_runtime plugin ref. | | `--codex-native` | no | Resume in Codex-native saved-login mode without provider API keys. | | `--model` | no | Override model name. | | `--model-provider` | no | Override model_provider plugin ref. | | `--strategy` | no | Override frontier strategy. Choices: `auto`, `mixed`, `explore`, `exploit`. | | `--cohort` | no | Cohort size override (exported as COHORT_SIZE). | | `--generations` | no | Maximum generations override (exported as MAX_GENERATIONS). | | `--server` | no | Disable --local mode (server mode). | | `--daemonize` | no | Use the same double-fork daemon launch path as praxist start. | | `--resume-policy` | no | Resume policy forwarded to praxist.run. Choices: `completed_generation`. Default: `completed_generation`. | | `--force` | no | Allow resume only when an old registry entry's process ownership cannot be verified. It never overrides a verified live controller. | | `--startup-timeout` | no | Seconds to wait for resume startup artifacts. Default: `30.0`. | | `--json` | no | Emit one JSON document on stdout instead of the operator summary. | ## `praxist setup` Pip-first Praxist host setup. Run this after installing the package and runtime extras. It writes only Praxist user-level configuration and Praxist-managed agent skill registrations; it does not install global agent CLIs or task-specific dependencies. ```text usage: praxist setup [-h] [--agent-system {claude_sdk,codex_sdk}] [--provider PROVIDER] [--model MODEL] [--api-key-stdin | --api-key-env API_KEY_ENV | --no-api-key] [--interactive] [--profile {codex-native,deepseek-api,openrouter-api,anthropic-api}] [--list-profiles] [--agent-managed] [--install-skills {codex,claude,none}] [--config-file CONFIG_FILE] [--json] [--dry-run] [--skip-doctor] ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | | `--agent-system` | no | Agent runtime selection to persist. Choices: `claude_sdk`, `codex_sdk`. | | `--provider` | no | Built-in provider name to configure. | | `--model` | no | Provider model name to persist. | | `--api-key-stdin` | no | Read the provider API key from stdin; a local TTY shows one * per character. | | `--api-key-env` | no | Read the provider API key from this environment variable. | | `--no-api-key` | no | Configure a supported no-key authentication route. | | `--interactive` | no | Review the License and User Agreement, choose optional privacy, and select a coherent runtime profile in a local TTY wizard. | | `--profile` | no | Apply one complete profile; a missing API key is requested in the local terminal. Choices: `codex-native`, `deepseek-api`, `openrouter-api`, `anthropic-api`. | | `--list-profiles` | no | List supported setup profiles as JSON and exit without changes. | | `--agent-managed`, `--codex-managed` | no | Report the read-only agent-managed first-use decision state and next required action as JSON. --codex-managed remains a compatibility alias. | | `--install-skills` | no | Install bundled skills for an agent host (interactive default: codex). Choices: `codex`, `claude`, `none`. | | `--config-file` | no | Config file to update (default: $PRAXIST_CONFIG_FILE or the user config). | | `--json` | no | Emit setup and readiness results as JSON. | | `--dry-run` | no | Validate and report changes without writing files. | | `--skip-doctor` | no | Skip the final readiness report. | ## `praxist start` Async launcher: starts ``python -m praxist.run run`` in a new session, redirects stdout/stderr to a run-local log file, and writes a registry entry under $PRAXIST_STATE_DIR/runs/. Pass --task-path / --model / --model-provider to override the resolved task and runtime configuration. ```text usage: praxist start [-h] [--task-path TASK_PATH] [--config-file CONFIG_FILE] [--agent-system {claude_sdk,codex_sdk}] [--runtime RUNTIME_REF] [--codex-native] [--run-dir RUN_DIR] [--resume] [--resume-from RESUME_FROM] [--resume-policy {completed_generation}] [--model MODEL] [--model-provider MODEL_PROVIDER_REF] [--strategy {auto,mixed,explore,exploit}] [--cohort COHORT] [--generations GENERATIONS] [--server] [--daemonize] [--startup-timeout STARTUP_TIMEOUT] [--json] ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | | `--task-path` | no | Task project directory (default: $TASK_PATH or the current directory). | | `--config-file` | no | Config file to load (default: $PRAXIST_CONFIG_FILE or the user config). | | `--agent-system` | no | Agent system the launched run will use. Default: $PRAXIST_AGENT_SYSTEM if set, else 'claude_sdk'. Recognised values: claude_sdk (default), codex_sdk. Choices: `claude_sdk`, `codex_sdk`. | | `--runtime` | no | Explicit ``agent_runtime:*`` plugin ref. Wins over the agent-system mapping when set. | | `--codex-native` | no | Use codex_sdk with native OpenAI and saved ChatGPT login, ignoring API-key and custom-endpoint settings from process/config/task env. | | `--run-dir` | no | Explicit run directory (default: /experiments/run__). | | `--resume` | no | Resume an existing run directory instead of requiring fresh artifacts. | | `--resume-from` | no | Path to an existing run directory to resume. Equivalent to --run-dir --resume. | | `--resume-policy` | no | Resume policy forwarded to praxist.run. Choices: `completed_generation`. Default: `completed_generation`. | | `--model` | no | Model name forwarded to the runtime; defaults depend on provider. | | `--model-provider` | no | Provider plugin ref (e.g. model_provider:deepseek_alias). Default cascades from agent system: claude_sdk → deepseek_alias when DEEPSEEK_API_KEY is set, then openrouter when OPENROUTER_API_KEY is set, then anthropic_messages; codex_sdk follows the same credential-aware selection and falls back to openai_compatible. | | `--strategy` | no | Frontier strategy (auto\|mixed\|explore\|exploit). Choices: `auto`, `mixed`, `explore`, `exploit`. Default: `auto`. | | `--cohort` | no | Cohort size override (exported as COHORT_SIZE). | | `--generations` | no | Maximum generations override (exported as MAX_GENERATIONS). | | `--server` | no | Disable --local mode (server mode). | | `--daemonize` | no | Double-fork the launcher before spawning so the workload survives when the launching shell's process tree is reaped. Required for sandboxed launcher contexts (agent tool shells, CI runners, Docker ``--init``). The default ``start_new_session=True`` path is fine for a normal terminal. | | `--startup-timeout` | no | Seconds to wait for startup artifacts before returning. A live run that exceeds the deadline remains in 'starting' state (default 30). Default: `30.0`. | | `--json` | no | Emit one JSON document on stdout instead of the operator table. | ## `praxist status` Merge the run registry written by ``praxist start`` with a cross-platform ``ps`` scan to list every Praxist run the operator should know about. Rows are tagged with their source: ``registry`` (managed run, PID alive), ``ps-only`` (matching process without a registry entry — e.g. started through a direct Python invocation), or ``stale`` (registry entry whose PID is gone). ```text usage: praxist status [-h] [--json] [--run-id RUN_ID] [--task-path TASK_PATH] [--active] [--latest] ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | | `--json` | no | Emit one JSON document on stdout instead of the plain-text table. | | `--run-id` | no | Show only this registry run id. | | `--task-path` | no | Show runs for this task directory. | | `--active` | no | Show only live local runs. | | `--latest` | no | Show only the newest matching run. | ## `praxist stop` ``praxist stop `` terminates one specific run via its registry entry. ``praxist stop --all`` terminates every Praxist-recognised process — by default the union of registry entries and ``ps``-scan matches. Registry-backed runs close new admission before discovery. Both modes send SIGTERM, wait --grace seconds, then SIGKILL any process still alive; registry-backed runs also perform a bounded stable-empty rescan for late children. ```text usage: praxist stop [-h] [--all] [--registry-only] [--ps-scan-only] [--grace GRACE_SECONDS] [--gc] [--dry-run] [--json] [run_id] ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | | `run_id` | no | Run id (filename stem of $PRAXIST_STATE_DIR/runs/.json). | | `--all` | no | Stop every recognised Praxist run (registry + ps-scan by default). | | `--registry-only` | no | With --all: only target registry-managed runs. | | `--ps-scan-only` | no | With --all: only target unregistered runs found by the process scan. | | `--grace` | no | Seconds to wait after SIGTERM before SIGKILL (default 5.0). Default: `5.0`. | | `--gc` | no | Remove stale registry entries. A stale entry is one whose recorded PID is no longer alive, or whose live command line no longer matches the prefix recorded at ``praxist start`` time (PID recycling). No signals are sent. | | `--dry-run` | no | Show what would be signalled without sending any signals. With --gc, list the would-be-removed entries without deleting any files. | | `--json` | no | Emit a JSON outcome document instead of the operator summary. | ## `praxist takeover` Open Codex or Claude Code and hand off a project to Praxist takeover. ```text usage: praxist takeover [-h] [--task-path TASK_PATH] [--codex-native | --configured-provider] [--operator {codex,claude}] [--yes] [--dry-run] [--json] ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | | `--task-path` | no | Research project to hand off (default: select locally or use the current directory). | | `--codex-native` | no | Use the no-key Codex-native takeover skill. | | `--configured-provider` | no | Use the configured-provider takeover skill. | | `--operator` | no | Agent CLI that hosts the takeover workflow. Default: codex. Choices: `codex`, `claude`. Default: `codex`. | | `--yes` | no | Launch without the final Enter confirmation. | | `--dry-run` | no | Show the redacted handoff without starting the agent CLI. | | `--json` | no | Emit the redacted handoff as JSON. | ## `praxist uninstall` Remove Praxist-managed CLI files, runtime environment, agent skills, configuration, state, and cache. Research projects, task environments, run directories, agent CLIs, Python, and uv are never removed. ```text usage: praxist uninstall [-h] [--venv-dir VENV_DIR] [--bin-dir BIN_DIR] [--skills-dir SKILLS_DIR] [--keep-user-data] [--dry-run] [--json] ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | | `--venv-dir` | no | Override the Praxist-managed virtualenv path. | | `--bin-dir` | no | Override the user bin directory containing Praxist entrypoints. | | `--skills-dir` | no | Override skill removal with one explicit managed directory. | | `--keep-user-data` | no | Keep Praxist configuration, registry state, product-usage state, and cache. | | `--dry-run` | no | Validate ownership and report removals without changing files. | | `--json` | no | Emit one machine-readable result document. | ## `praxist user-agreement` Review the Praxist License and User Agreement or inspect acceptance status. ```text usage: praxist user-agreement [-h] ... ``` | Argument | Required | Description | |---|---:|---| | `-h`, `--help` | no | show this help message and exit | --- # Skills Reference This catalog is generated from bundled `SKILL.md` front matter. Each skill file is the sole definition of its activation contract and workflow. Mechanism abbreviations use the compact definitions in the [Glossary](https://praxist.sapient.inc/en/docs/about/glossary). | Skill | Invoke in Codex | Purpose | |---|---|---| | `praxist-control` | `$praxist-control` | Start, stop, resume, inspect status, open the independent read-only foreground TUI, detect active runs, and safely repair Praxist run lifecycle boundaries from a supported agent interface. Use when the user asks to launch a Praxist task, stop or kill a run, continue/resume an interrupted run, restart from the latest safe generation, query current run progress, open or exit a live monitor, detect or list currently running Praxist tasks in the environment, inspect generation status, view incubator/frontier/leaderboard performance, check hardware load, handle interrupted PI panel or Gems reset boundaries, inspect whether a task directory is runnable, or control Praxist lifecycle commands with `praxist start`, `praxist stop`, `praxist status`, `praxist --monitor`, `praxist resume`, or `praxist resolve`.
Source: `skills/praxist-control/SKILL.md` | | `praxist-diagnostic` | `$praxist-diagnostic` | Diagnose Praxist run health, artifact integrity, research-loop completeness, generation-scoped DIG/QD, PI/Gems/frontier/incubator consistency, peer memory freshness, diversity HHI, hardware utilization, LLM/runtime friction, task harness health, sustained low-performance causes, strongest variants/Pareto front, strong-variant lineage, and human-readable run reports. Use when the user asks an agent to investigate whether a current or historical Praxist run is healthy, why progress or performance is weak, whether artifacts or promotions are missing, whether guard or resource issues are blocking peers, to produce a detailed agent behavior analysis report, to generate an A/B/C run report, or to improve or optimize a task after diagnosis using task-directory-only parameter and prompt changes. Default diagnostics are analysis-only; explicit improvement mode may stop the selected run and edit task-level configuration/prompts, but must not modify Praxist core logic.
Source: `skills/praxist-diagnostic/SKILL.md` | | `praxist-interactive-task-init` | `$praxist-interactive-task-init` | Build a Praxist task project through a confirmation-first interactive agent workflow. Use when a user wants Praxist task initialization with human confirmation of research goals, constraints, metrics, ranking rules, evaluation protocol, compute budget, baseline handling, or launch readiness; when the user asks for an interactive task init skill; or when the agent should propose a task harness first and ask the user to approve or revise it before writing files.
Source: `skills/praxist-interactive-task-init/SKILL.md` | | `praxist-onboarding` | `$praxist-onboarding` | Establish detailed context for Praxist before helping a user install, configure, operate, troubleshoot, or extend the system. Use when the user is new to Praxist, has installed or is about to install the package, asks what Praxist does, asks how Praxist works, asks about `praxist` commands, task projects, runs, peers, generations, frontier/incubator/Gems, PI panels, configuration, API keys, model providers, agent runtimes, plugin architecture, run artifacts, software boundaries, or asks an agent to inspect whether the local environment is ready for Praxist. Do not use for a specific research task's domain science unless Praxist system context is needed first.
Source: `skills/praxist-onboarding/SKILL.md` | | `praxist-runtime-install` | `$praxist-runtime-install` | Install Praxist runtime dependencies and configure user-level Praxist provider credentials for a source checkout or pip-installed environment. Use when the user asks an agent to install or repair Praxist requirements, prepare a Praxist host, create the Praxist Python environment, install the `praxist` CLI, install Claude SDK or official Codex SDK runtime extras, install source-checkout test/dev dependencies, persist API keys or provider settings, verify imports, or diagnose missing dependencies. For source checkouts include the repository test/dev dependency group; for pip package installs keep the install runtime-only. Do not use for docs, task-specific training, dataset, benchmark, or experiment dependencies.
Source: `skills/praxist-runtime-install/SKILL.md` | | `praxist-scientific-research` | `$praxist-scientific-research` | Gather task-agnostic scientific research context for Praxist task projects using no-key public literature/database/open-access lookup, agent-host web search when available, local project documents, and source/provenance notes. Use when an agent needs to identify domain metrics, benchmarks, prior art, scientific databases, open-access provenance, high-value research directions, or literature-backed hypotheses for a Praxist task without starting a run, changing Praxist core logic, or treating literature as measured task performance.
Source: `skills/praxist-scientific-research/SKILL.md` | | `praxist-takeover` | `$praxist-takeover` | Orchestrate first-use Praxist onboarding, current task initialization, validation, and detached run launch from a supported agent interface. Use when a user wants a one-command or one-conversation path from an existing runnable research project to a started Praxist run, asks to onboard plus initialize plus start, asks for repo-to-task-to-run setup, or wants the agent to prepare and launch Praxist with minimal interaction. This skill composes onboarding, full task initialization or task-harness repair, runtime checks, control start, and confirmation gates without changing Praxist core.
Source: `skills/praxist-takeover/SKILL.md` | | `praxist-takeover-codex` | `$praxist-takeover-codex` | Onboard, initialize or repair, validate, and launch a Praxist research task in Codex-native mode through the official Codex SDK runtime and the operator's existing saved ChatGPT login, with catalog-verified gpt-5.6-luna as the default model and without requesting, storing, or using an API key. Use when a user wants a low-interaction no-key Praxist takeover, explicitly requests Codex-native mode, has no provider key, or invokes this skill with no additional text after already logging in to Codex.
Source: `skills/praxist-takeover-codex/SKILL.md` | | `praxist-task-initialization` | `$praxist-task-initialization` | Convert an existing runnable computer-based research project into a formal Praxist task project, or repair a task harness that fails task-init validation. Use when a user wants an agent to transform AI algorithm, robotics, control, simulation, SLAM, LLM, optimization, or other executable research code into a Praxist task directory with task.yaml, baseline harness, evaluator, baseline performance records, robust metric/ranking policy, protocol-integrity checks, reachable task-justified durable/Pareto retention lanes, role prompts, audit rules, dataset/simulator metadata, high-value research directions, initial-generation DIG plus independently controlled QD, continuous-evolution/Gems research-loop settings, run-report tooling, and hardware-aware or user-selected fixed Praxist run parameters. The skill requires a project that already runs on the current machine or in an available environment/container. Abort when required code, data/simulator assets, or declared runtime dependencies are missing.
Source: `skills/praxist-task-initialization/SKILL.md` | | `terminal-line-plot` | `$terminal-line-plot` | Draw readable ASCII line charts directly in the terminal from numeric series, command output, CSV, JSON, or manually extracted points. Use when the user asks for a curve, trend line, score progression, leaderboard trend, metric history, or any plot that should be visible in a CLI/chat transcript without opening a GUI or writing image files.
Source: `skills/terminal-line-plot/SKILL.md` | --- # Core API Reference This page is generated from public docstrings in `praxist.core`. ## Registry ::: praxist.core.registry ## Task Projects ::: praxist.core.task_project ## Protocol ::: praxist.core.protocol ## Runtimes ::: praxist.core.runtimes ## Modeling ::: praxist.core.modeling ## Budget ::: praxist.core.budget ## Credentials ::: praxist.core.credentials ## Prompt Layout ::: praxist.core.prompt_layout ## Tool Servers ::: praxist.core.tool_servers ## Role Skills ::: praxist.core.role_skills ## Workflow ::: praxist.core.workflow ## Storage ::: praxist.core.storage ## Trajectory ::: praxist.core.trajectory ## Replay ::: praxist.core.replay ## Execution Guards ::: praxist.core.execution_guards ## Source Snapshot ::: praxist.core.source_snapshot --- # Plugin API Reference This page documents executable generic plugin boundaries. ## Agent Runtimes ::: praxist.plugins.agent_runtimes.claude_sdk.adapter ::: praxist.plugins.agent_runtimes.codex_sdk.adapter ## API Providers ::: praxist.plugins.model_providers.openrouter.adapter ::: praxist.plugins.model_providers.anthropic_messages.adapter ::: praxist.plugins.model_providers.openai_compatible.adapter ::: praxist.plugins.model_providers.deepseek_alias.adapter ## Workflow Stage ::: praxist.plugins.workflow_stages.research_loop.startup ::: praxist.plugins.workflow_stages.research_loop.stage ::: praxist.plugins.workflow_stages.research_loop.c5_materializer ::: praxist.plugins.workflow_stages.reviewer_stub.adapter ## Tools ::: praxist.plugins.tools.evaluation_tools.adapter ::: praxist.plugins.tools.finding_graph_query.adapter ::: praxist.plugins.tools.frontier_tools.adapter ::: praxist.plugins.tools.memory_tools.adapter ::: praxist.plugins.tools.prior_work_tools.adapter ::: praxist.plugins.tools.literature_lookup.adapter ## Graph Maintainer ::: praxist.plugins.graph_maintainers.finding_graph_mvp.adapter ::: praxist.plugins.graph_maintainers.finding_graph_mvp.engine ::: praxist.plugins.graph_maintainers.finding_graph_mvp.cli ## Budget Policy ::: praxist.plugins.budget_policies.default_basic.policy --- # CLI and Operator API Reference This page documents Python operator entrypoints and support scripts. ## Run CLI ::: praxist.run ## Deliverables ::: praxist.deliver ::: scripts.deliver_auto_research ## Task Spec Compatibility ::: praxist.task_spec ## Fake Workflow Fixture ::: praxist.testing.fake_workflow_fixture --- # Task Template API Reference Tracked templates demonstrate the task-project layout. They are not a production task catalog or complete worked examples. Use [Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates) to choose the right asset, and inspect [Rocket Booster Recovery](https://praxist.sapient.inc/en/docs/examples/rocket-booster-recovery) for a complete runnable integration. ## Template Runner ::: templates.tasks.template.runner ## Toy Math Runner ::: templates.tasks.toy_math.runner ## SAM Optimizer Reference Evaluation ::: templates.tasks.sam_optimizer.evaluations.pareto_tiered.evaluator ## SAM Optimizer Reference Evaluation Runner ::: templates.tasks.sam_optimizer.evaluations.pareto_tiered.run ## Template Evaluation ::: templates.tasks.template.evaluations.pareto_tiered.evaluator ## Template Evaluation Runner ::: templates.tasks.template.evaluations.pareto_tiered.run ## Toy Math Evaluation ::: templates.tasks.toy_math.evaluations.pareto_tiered.evaluator ## Toy Math Evaluation Runner ::: templates.tasks.toy_math.evaluations.pareto_tiered.run ## Configuration Source The executable API objects above are the purpose of this reference page. Configuration, evidence maturity, Deep Innovation Gate (DIG), Quality-Diversity (QD), Gems, lane routing, artifact ownership, and scheduler requirements for templates are defined once in [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects). The tracked task descriptors under `templates/tasks/` are validated fixtures of that contract, not a second schema definition. --- # Glossary This glossary gives short navigation definitions. The linked pages own the full contracts. | Term | Meaning | Canonical detail | |---|---|---| | Task project | External runnable research problem supplied to Praxist | [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects) | | Template | Replaceable task-project scaffold or deterministic smoke fixture | [Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates) | | Example | Complete runnable reference project with task-owned code, evidence, and tests | [Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates) | | Peer | One research agent working within a generation | [Research Loop](https://praxist.sapient.inc/en/docs/guides/research-loop-variant-generation-flow) | | Generation | A cohort of peer work followed by a research-planning boundary | [Research Loop](https://praxist.sapient.inc/en/docs/guides/research-loop-variant-generation-flow) | | Finding | Structured report of observed evidence or a reusable research lesson | [Architecture](https://praxist.sapient.inc/en/docs/concepts/architecture) | | Validation signal | Compact, non-durable evidence retained for validation, repair, or diagnostic follow-up | [Architecture](https://praxist.sapient.inc/en/docs/concepts/architecture) | | Incubator | Task-defined durable lower-admission library for complete, credible candidates | [Flexibility Controls](https://praxist.sapient.inc/en/docs/guides/research-loop-flexibility-controls) | | Frontier | Durable task-defined promoted evidence used by planning and reporting | [Architecture](https://praxist.sapient.inc/en/docs/concepts/architecture) | | Gems | Compact selected research memory used when a task enables periodic reset | [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects) | | Principal Investigator (PI) | Independent planning agent that proposes next-generation work from committed evidence | [Panel Topology Prompts](https://praxist.sapient.inc/en/docs/concepts/panel_topology_prompts) | | Chair | Planning agent that compares PI proposals and commits one coherent agenda | [Panel Topology Prompts](https://praxist.sapient.inc/en/docs/concepts/panel_topology_prompts) | | Deep Innovation Gate (DIG) | Deep-reasoning innovation process that compares candidate mechanisms before implementation | [Deep Innovation Gate](https://praxist.sapient.inc/en/docs/guides/deep-innovation-gate) | | Quality-Diversity (QD) | Allocation principle that preserves varied strong candidate plans | [Quality-Diversity Allocation](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator) | | Herfindahl-Hirschman Index (HHI) | Concentration measure used to diagnose whether planned or realized work collapsed into too few categories | [Quality-Diversity Allocation](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator) | | Task harness | Task-owned evaluator, baseline, prompts, roles, and evidence contracts | [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects) | | Codex-native mode | Explicit Codex SDK route authenticated by a saved ChatGPT login | [Quickstart](https://praxist.sapient.inc/en/docs/getting-started/quickstart) | --- # Praxist User Agreement - **Agreement version:** 2026-08-28 - **Effective date:** 28 August 2026 ## Chapter 1: General Provisions ### 1.1 Scope This Agreement is a complete and legally binding contract between the user (the **User**) and Sapient Intelligence Pte Ltd (the **Company**), concerning the Praxist software and related services (the **Services**). The Services may include the locally installed Praxist software, command-line interfaces, packaged agent skills, documentation, updates, support, and optional Company-operated network services made available with a Praxist release. ### 1.2 Acceptance The User accepts this Agreement by selecting **I agree and continue** in the Praxist first-use experience, by otherwise recording an explicit acceptance, or by actually using the Services. The User confirms that they have had an opportunity to review the complete Agreement, including Appendix A and any supplementary terms incorporated by reference. A User who does not agree must cancel setup and stop using the Services. Praxist stores a minimal acceptance record for the current operating-system user. It contains the Agreement version and digest, acceptance time, and whether the choice was made directly or relayed by an Agent. This record stays on the local machine and is not included in optional product-usage events. ### 1.3 Governing framework This Agreement is formulated in accordance with the laws of the Republic of Singapore and other applicable laws and regulations. Its purpose is to define the rights and obligations of both parties and protect their legitimate rights and interests. ### 1.4 Revisions The Company may revise this Agreement in response to business development, changes in applicable law, or changes to the Services. The current Agreement will be identified by version and published with the relevant release and in the official Praxist documentation. Where renewed acceptance is required, Praxist will present the revised Agreement before continuing first-use setup. Continued use after the applicable notice or acceptance process constitutes acceptance of the revised Agreement. A User who refuses a revision must stop using the affected Services. ## Chapter 2: Eligibility, Installation, and Credential Management ### 2.1 Eligibility Users must be natural persons with full civil capacity, legal persons, or other organizations legally capable of entering this Agreement. A minor may use the Services only with the consent and supervision of a legal guardian, who bears responsibility for that use to the extent required by law. ### 2.2 Installation and configuration The current Praxist software does not require a separate Praxist account for local installation or local research operation. The User is responsible for providing accurate configuration, selecting an authorized model-provider or runtime account, and ensuring that the research project and its dependencies may lawfully be used. If a future Company-operated service requires registration, the User must provide true, accurate, and complete registration information and must not impersonate another person, create accounts in bulk, or otherwise abuse registration. ### 2.3 Third-party accounts and credentials Accounts, subscriptions, API keys, and saved login credentials used by Praxist may be issued by third-party providers. Their ownership and permitted use are governed by the User's agreement with the relevant provider. Praxist stores provider configuration locally when the User asks it to do so and does not claim ownership of the User's third-party account. The User shall not transfer, rent, lend, sell, or misuse any Company-operated account if such an account is provided separately. ### 2.4 Security The User shall safeguard passwords, API keys, saved login credentials, and other authentication material and is responsible for operations performed with those credentials. The User shall promptly notify the Company and the relevant third-party provider of suspected unauthorized access. The Company may provide reasonable assistance but is not liable for loss caused by the User's failure to protect credentials, except where applicable law provides otherwise. ## Chapter 3: Service Content and Usage Code of Conduct ### 3.1 Service content Praxist provides a task-agnostic framework for measurable, computer-executable research. It can coordinate agent runtimes, task-owned experiments, evidence, research planning, and reporting. Unless expressly stated otherwise, Praxist does not supply the User's research project, datasets, simulators, evaluator, task-specific dependencies, scientific acceptance criteria, computing resources, or third-party model service. The specific Services available to a User are those included in the installed release or displayed in official documentation. The Company may add, remove, or improve Services and will publish material changes through official release materials or documentation. ### 3.2 Usage rules The User shall comply with applicable law, public order, and good morals and shall not use the Services to: 1. create or disseminate content prohibited by applicable law, including unlawful political propaganda, threats to national security, illegal gambling, malicious hacking tools, obscene content, or unlawful discriminatory content; 2. infringe intellectual property, portrait, reputation, privacy, or other legitimate rights, impersonate another person, or disclose another person's private information without authority; 3. commit fraud, extortion, harassment, deception, or other illegal acts; 4. maliciously attack, crack, disrupt, tamper with, or steal data from Praxist or any connected service; 5. use a Service outside the scope authorized by the Company or a relevant third-party provider; or 6. violate applicable national or regional laws, regulations, sanctions, or published usage rules. ### 3.3 Service restrictions The Company may limit usage of Company-operated online Services based on service capacity, security, legal requirements, or abusive use. Third-party model providers and infrastructure providers may impose their own quotas and usage limits. This clause does not give the Company remote control over the User's lawful local computing resources or task project. The Company may suspend or terminate access to a Company-operated Service for material breach without compensation except where applicable law requires otherwise. ## Chapter 4: Intellectual Property Rights ### 4.1 Praxist materials The Company retains rights in Company-authored software, trademarks, patents, algorithms, service designs, interface designs, documentation, and written content to the fullest extent permitted by law. Use and redistribution of software or documentation are also subject to the license terms accompanying the relevant distribution. Third-party software, models, datasets, and other materials remain subject to their respective owners' rights and licenses. Nothing in this Agreement overrides an applicable open-source or third-party license. ### 4.2 User-generated content Unless the parties enter a separate written agreement, the User retains the copyright and other rights they hold in research inputs, task projects, and content generated through the Services. The User remains responsible for any third-party terms that apply to a model, dataset, runtime, or other component used to produce that content. ### 4.3 Local inputs and voluntary submissions Local processing by Praxist does not transfer the User's rights in task content, prompts, research results, files, or project paths to the Company. Those materials are not included in optional product-usage events described in Appendix A. If the User separately and voluntarily submits material to the Company for support, feedback, or another requested service, the User grants a non-exclusive license limited to providing that service and improving Praxist, subject to applicable privacy duties and any separate written terms. ### 4.4 User warranty The User warrants that submitted or processed content is within the User's right to use and does not infringe third-party rights. The User shall bear liability for disputes and losses caused by content the User had no right to use, including legally recoverable damages, litigation costs, and attorney fees. ## Chapter 5: Suspension, Termination, and Modification of Services ### 5.1 Availability Company-operated online Services may be suspended for maintenance, upgrades, malfunctions, force majeure, security incidents, or legal requirements. The Company will provide reasonable notice where practicable and will announce restoration when appropriate. Locally installed Praxist software may remain available, but its operation can depend on third-party runtimes, providers, networks, or infrastructure outside the Company's control. ### 5.2 User breach If the User materially breaches this Agreement, uses a Company-operated Service unlawfully, or provides false information where registration is required, the Company may restrict or terminate access to that Service and any associated Company-operated account. This does not authorize the Company to delete the User's local research project. Local and hosted data, if any, will be handled under the applicable documentation and law. ### 5.3 Service changes The Company may modify or terminate part or all of the Services as business or legal requirements change. Material changes will be published through official release materials, documentation, or service notices and take effect after any required notice period. Where a paid Company-operated Service is terminated, the Company will handle outstanding matters according to the applicable order terms and law. ## Chapter 6: Rights and Obligations of Both Parties ### 6.1 Company rights and obligations The Company may: 1. provide the Services under this Agreement, manage Company-operated Services, and respond to misuse; 2. revise this Agreement and official rules in accordance with Section 1.4; 3. use reasonable efforts to maintain Company-operated systems, publish software updates, and respond to reasonable feedback; 4. protect personal information in accordance with applicable law and avoid unauthorized disclosure or misuse; 5. refrain from using the Services to conduct illegal activity or infringe the legitimate rights of Users or third parties; and 6. collect the bounded product-usage data in Appendix A only after a separate, explicit opt-in. Acceptance of this User Agreement alone does not enable product-usage collection. ### 6.2 User rights and obligations The User may and shall: 1. use the Services under this Agreement and submit reasonable suggestions; 2. exercise applicable rights to inquire about, correct, or delete personal information and to close any Company-operated account; 3. comply with this Agreement and refrain from illegal acts or conduct that harms the Company or third parties; 4. safeguard credentials and accept responsibility for authorized operations performed with them; 5. report material faults or vulnerabilities responsibly and refrain from malicious exploitation; and 6. comply with relevant terms when using third-party providers or services through Praxist. ## Chapter 7: Liability for Breach of Contract ### 7.1 User breach If the User breaches this Agreement, the Company may suspend or restrict Company-operated Services or terminate a Company-operated account. The User shall compensate the Company for losses recoverable under applicable law, including direct losses and reasonable litigation and attorney fees. ### 7.2 Company breach If the Company breaches this Agreement by failing to perform an applicable service obligation or infringing the User's legitimate rights, the Company shall bear liability required by law. To the extent permitted by law, the Company is not liable for indirect losses or lost anticipated profits, and its aggregate liability shall not exceed the service fees the User actually paid to the Company for the affected Service. ### 7.3 Force majeure Neither party is liable for a failure caused by force majeure, including earthquakes, floods, typhoons, war, policy changes, widespread system failure, or cyberattack, to the extent recognized by law. The affected party shall give prompt notice where practicable and take reasonable steps to reduce loss. ## Chapter 8: Dispute Resolution ### 8.1 Governing law The formation, performance, interpretation, and dispute resolution of this Agreement are governed by the laws of Singapore. ### 8.2 Arbitration The parties shall first attempt to resolve any dispute arising out of or in connection with this Agreement through good-faith negotiation. If negotiation fails, either party may submit the dispute to the Singapore International Arbitration Centre (SIAC) for arbitration under its rules then in force. ## Chapter 9: Miscellaneous Provisions ### 9.1 Severability If any clause is held invalid or unenforceable, the remaining clauses remain in force to the extent permitted by law. ### 9.2 Supplementary terms Matters not covered here may be governed by a separate written supplementary agreement, which will have the same legal effect when validly entered into by the parties. ### 9.3 Incorporated documents Appendix A, applicable service-specific terms, and official notices expressly incorporated into this Agreement form part of it. Product documentation that only explains operation does not silently expand the data-collection scope in Appendix A. ### 9.4 Term and interpretation This Agreement takes effect when the User records acceptance or begins using the Services and remains effective until use and any Company-operated account or Service relationship end, subject to clauses that survive by their nature. The Company may interpret and revise this Agreement subject to applicable law and Section 1.4. - **Company / service operator:** Sapient Intelligence Pte Ltd - **Contact:** praxist@sapient.inc - **Release date:** 28 August 2026 Appendix A is the separately authored [Praxist User Data Collection Notice](https://praxist.sapient.inc/en/docs/legal/product-usage-data-notice). It is incorporated into this Agreement, but product-usage collection remains disabled unless the User gives the separate opt-in described there. --- # Praxist Privacy Notice (Product Usage Data) **Last updated:** 27 August 2026 **Related documents:** [Fair Source License Agreement (Version 1.0)](https://github.com/sapientinc/praxist/blob/main/LICENSE.md); [Praxist User Data Collection Notice](https://praxist.sapient.inc/en/docs/legal/product-usage-data-notice) (Notice version 3); [Product Usage Technical Documentation](https://praxist.sapient.inc/en/docs/operations/DOCUMENTATION) --- ## 1. Overview This Notice applies only to Praxist's **optional product-usage data collection**. The data controller is Sapient Intelligence Pte Ltd ("we", "us"). Core principles: - **Voluntary.** Collection is predicated on your explicit consent and is never mandatory. - **Revocable.** You may withdraw consent at any time; collection stops immediately upon withdrawal. - **No impact on use.** Declining or withdrawing consent does not affect the installation of Praxist or any research functionality. - **Pseudonymized and minimized.** We collect only pseudonymized lifecycle data within a closed schema — never any task content, input data, or research results. ## 2. Preconditions for Collection (Voluntary Consent Mechanism) Praxist collects product-usage data only when **all three** of the following conditions are met: 1. the installed build contains an approved collection transport; 2. product-usage collection is enabled for that build; and 3. you separately and explicitly select **"Share product usage"** after this Notice is made available for review, or reply `Yes` / `Agree` to consent prompt shown during installation. Please note: - Accepting the Praxist User Agreement or the Fair Source License Agreement does **not** constitute consent to data collection; the two are independent; - when no choice has been made (status "unset"), **nothing is collected or uploaded**; - in non-interactive environments (no terminal input), no consent prompt is shown and collection remains off; - the same rule applies to Agent-assisted installation: only `Yes` / `Agree` (consent) and `No` / `Disagree` (refusal) are recognized; any other wording is treated as no consent given; - this Notice is versioned (currently Notice version 3). When the Notice changes, its version number increases; consent you previously recorded does not carry over to a new version and will be requested again. ## 3. What We May Collect (If You Consent) The product-usage protocol uses a **closed schema** (Schema V2): only the fields below are permitted, and any additional field is rejected by both the client and the server. ### 3.1 Common lifecycle fields | Field | Description | | --- | --- | | `schema_version` | Version of the product-usage event structure (currently 2) | | `praxist_version` | Public Praxist version | | `consent_notice_version` | The Notice version you consented to | | `environment_id` | Environment identifier: a random UUID generated locally, stable across Research Runs within one Praxist environment (see Section 5) | | `telemetry_run_id` | Independent random identifier for a single Research Run | | `event_id` | Random identifier for a single event, used for correlation and deduplication | | `event_sequence` | Sequence number within the same Research Run | | `event_type` | Lifecycle event type (see Section 3.2) | | `occurred_at` | Client-side event time (UTC, to the second) | | `error_summaries` / `error_summaries_truncated` | Bounded structured error-category counts, and whether the bounded list was truncated (see Section 3.3) | ### 3.2 The four lifecycle events | Event | What it records | | --- | --- | | `run_started` | The generation ordinal, the planned Peer count, and aggregate counts of Peers in the planning, running, completed, cancelled, failed, and unknown states at the run-start boundary | | `generation_finished` | The same fields recorded at a durable generation boundary (state counts sum to the planned Peer count; "completed" means only that a Peer lifecycle returned normally — it does not assert that any scientific result is valid) | | `run_finished` | Active run duration in complete minutes (capped at 43,200 minutes, i.e. 30 days) and whether the cap was reached | | `run_reconciled` | The same duration fields, recorded when a previously unfinished run is resumed and trustworthy terminal processing is later completed | ### 3.3 Structured error summaries Each lifecycle event may carry at most 16 grouped error summaries. Each group may contain only the following closed values: - `scope`: run / generation / peer - `stage`: setup / launch / execution / finalization / reconciliation - `error_type`: configuration / resource / orchestration / runtime / external_dependency / storage / unknown - `error_code`: PRX-CAPACITY / PRX-PEER-LAUNCH / PRX-PEER-RUNTIME / PRX-RUNTIME / PRX-RUN-FAILED / PRX-UNKNOWN - `reason_code`: auth_error / quota_exhausted / rate_limited / timeout / provider_unavailable / runtime_error / tool_unavailable / invalid_request / budget_denied / budget_expired / capacity_unavailable / process_start_failed / state_unreadable / unexpected_termination / unknown - `count`: number of matching errors (capped at 65,535), plus whether the cap was reached This structure is **technically incapable** of containing raw error messages, logs, stack traces, provider responses, or arbitrary text. ### 3.4 Time fields - `occurred_at` is generated by the local Praxist client at a lifecycle milestone, converted to UTC using the local system clock (and may therefore reflect clock inaccuracy); - `received_at` is added by the server-side Collector after validation and cannot be supplied or altered by the client. It represents arrival time and is used only for storage management and retention calculation. ## 4. What We Do Not Collect Product-usage events do **not** include: - research task content, prompts, research results, files, filenames, project paths, or commands; - environment variables, API keys, saved login credentials, logs, stack traces, raw error messages, or arbitrary text; - model names, service-provider names, provider responses, or account information; - names, email addresses, operating-system details, hardware information, Python version, client time zone, cookies, or arbitrary request headers; - individual Peer identities or individual Peer outputs; - IP addresses (event bodies never contain them; the temporary handling of network connection information is described in Section 6). ## 5. Pseudonymization Methodology Praxist uses **pseudonymization**, not full anonymization: - **Generation.** `environment_id`, `telemetry_run_id`, and `event_id` are all randomly generated via standard UUIDv4 (`uuid4()`); - **No derivation.** These identifiers are not derived from usernames, accounts, device serial numbers, MAC addresses, IP addresses, hostnames, project paths, or task content, and they are not simple incrementing sequences; - **Persistence.** `environment_id` is generated the first time an environment needs one and stored locally in `environment.json` (readable and writable only by the current user, permission 0600); it remains stable across runs within that environment. `telemetry_run_id` and `event_id` are generated per run and per event; - **Local path hashing.** Local run-state files are named with the SHA-256 hash of the run path; these files **stay on your machine and are never uploaded**; - **Honest characterization.** Because `environment_id` is stable across runs, events from the same environment could in theory be linked to one another — this is pseudonymized data, not fully anonymous data. However, it cannot directly identify you or your device, and we commit to **never using it in any way** to identify an individual (consistent with our commitment in Section 7). ## 6. Network and Transmission - **Production endpoint:** `https://telemetry.theaiscientist.com/v1/events` (HTTPS encryption, server certificate verification, no redirect following); - Requests carry only a fixed, protocol-level User-Agent (`Praxist-Product-Usage/2`); each request is capped at 32 KB and at most 50 events per batch; network timeout is 2 seconds; - Transmission failures never block or affect Research Runs; events that fail to send are stored locally and delivered later when the network is available (see Section 8); - **Server-side handling of connection information.** Product-usage event bodies contain no IP addresses, cookies, or arbitrary request headers. Network services necessarily process connection information transiently to deliver requests, protect the service, and apply rate limits. The Collector deployment disables access logging and strips forwarded IP, cookie, and client User-Agent headers before application processing; none of them are persisted as product-usage event data. ## 7. Purposes of Use Collected product-usage data is used **only** for: 1. analyzing product reliability; 2. improving Praxist's features, performance, and user experience; and 3. anonymous, aggregate-level statistics that are never presented in a way that could identify a particular user. We commit that we will **not**: - sell product-usage data to any third party; - use it for advertising or marketing; or - link it with any other dataset about you or your users in order to identify an individual. ## 8. Storage, Retention, and Deletion - **Server-side storage.** Managed PostgreSQL, accessed over a private endpoint; no public internet database exposure; - **Retention.** Delivered raw events are retained for **at most 180 days** from server receipt and are then deleted by a scheduled retention process (the job runs at least once daily; by design, events enter the deletion window on day 179, leaving one day of scheduling slack so the stated 180-day ceiling is never exceeded); - **Local unsent events.** Stored in the local SQLite outbound queue (`outbox.sqlite3`, readable and writable only by the current user), cleared after successful delivery, and deleted immediately upon withdrawal; - **Honest note.** The server does not offer an interface to delete **already-delivered** events by environment identifier. Withdrawal stops future collection and deletes local unsent events; delivered events are deleted automatically when the retention period expires. ## 9. Your Rights and How to Exercise Them | Action | Command | | --- | --- | | View the complete Notice text | `praxist product-usage notice` | | Check current consent status | `praxist product-usage status --json` | | Record consent | `praxist product-usage consent` (or select "Share product usage" during first use) | | **Withdraw consent at any time** | `praxist product-usage withdraw` | - Withdrawal immediately stops all future capture and deletes all local unsent events; - Declining or withdrawing **does not affect** the installation, operation, or any research functionality of Praxist; - You may also exercise your rights of access, rectification, deletion, or complaint through the contact point in Section 13. ## 10. Data Recipient and Cross-Border Arrangements - **Data recipient:** Sapient Intelligence Pte Ltd (the same legal entity as the Licensor under the license agreement) - **Server location:** Johor, Malaysia - **Cross-border transfer:** If you are located in mainland China and choose to opt in, your consent constitutes authorization for the transfer of the above data — which contains no names, contact details, accounts, IP addresses, or content data (see the closed lists in Sections 3 and 4) — to the Collector in Malaysia, limited to the fields enumerated in the Section 3 closed schema. Users in the EU/EEA are covered by Section 10.1. ### 10.1 Supplementary Notice for Users in the European Union and EEA (GDPR) Praxist is distributed globally, and users in the EU/EEA may likewise choose to opt in to product-usage collection. Under the GDPR, pseudonymized data remains personal data, and we apply the following rules to such data: - **Legal basis.** Your explicit consent only (GDPR Art. 6(1)(a)) — data collection is off by default and is enabled only when you actively select "Share product usage"; - **Withdrawal.** You may withdraw consent at any time via `praxist product-usage withdraw`; withdrawal does not affect the lawfulness of processing carried out on the basis of consent before its withdrawal (Art. 9(3)); - **Cross-border transfer.** Data is transferred to and processed by the Collector in Johor, Malaysia. This transfer relies on the explicit consent you give at opt-in (the derogation in GDPR Art. 49(1)(a)) and is limited to the fields enumerated in the Section 3 closed schema of this Notice; - **Your rights.** Access, rectification, restriction of processing, data portability, objection, and erasure. Send access, rectification, or deletion requests to praxist@sapient.inc; - **Honest note on erasure.** As described in Section 8, the server currently is not able to search or delete **already-delivered** events by environment identifier; delivered events are deleted automatically at most 180 days after receipt. You may include the `environment_id` from your local `environment.json` in a deletion request to verify ownership of the environment, and we will manually process your request after verification; - **We do not:** sell personal data, use it for advertising or marketing, or carry out automated decision-making with legal or similarly significant effects; - **Supervisory authority.** You have the right to lodge a complaint with the data protection supervisory authority of your member state. ## 11. Data Security - Production traffic is HTTPS-encrypted end to end with certificate verification, and redirects are refused; - Closed schema: both the client and the server reject any field outside the schema; - The server enforces global and per-client rate limits and a storage capacity ceiling; a master ingestion switch can shut down the collection entry entirely in an emergency; - Local consent records, the environment identifier, and the outbound queue are readable and writable only by the current user (0600/0700 permissions); - Events are deleted automatically when the retention period expires — no action required from you; - Security contact: praxist@sapient.inc ## 12. Changes to This Notice We may revise this Notice and the in-product notice as the product evolves. When the notice content changes, the Notice version number increases; consent you recorded is valid only for the version it was given against, and renewed consent will be requested at first use after an upgrade. If this Notice and the in-product notice diverge, the version presented to you at the time of consent prevails. ## 13. Contact Us - Privacy contact: praxist@sapient.inc - Company: Sapient Intelligence Pte Ltd (incorporated in Singapore) - Commercial licensing and license-agreement matters: see the contact details at the end of the Fair Source License Agreement (Version 1.0) (praxist@sapient.inc) --- # Appendix A: Praxist User Data Collection Notice **Notice version:** 3 **Effective date:** 19 August 2026 ## 1. Separate and Optional Consent The User may help improve Praxist by sharing pseudonymized (pseudonymous) product-usage data. Praxist collects it only when all of the following are true: 1. the installed build contains an approved collection transport; 2. product-usage collection is enabled for that build; and 3. the User separately selects **Share product usage** or otherwise provides an explicit supported opt-in after this Notice is made available for review. Accepting the Praxist Fair Source License and User Agreement does **not** provide this optional consent. If the installed build has no approved collection capability, or if consent is unset or denied, no product-usage events are collected or sent. Research operation is not reduced when the User declines. A temporary network or collector outage after opt-in may leave bounded events in the local outbox for later delivery; it does not affect the Research Run. Development builds send authorized events to the fixed Praxist development collector at `http://45.78.201.249/v1/events`. This development transport is plain HTTP and is intended only for internal development or test data. Formal releases send authorized events to the Praxist production collector at `https://telemetry.theaiscientist.com/v1/events` using HTTPS certificate verification and without following redirects. ## 2. Data Praxist May Collect After Opt-In The product-usage protocol has a closed schema. It may collect only the fields described below. ### 2.1 Common lifecycle fields - `schema_version`: version of the product-usage event structure; - `praxist_version`: public Praxist version; - `consent_notice_version`: version of this Notice; - `environment_id` (Environment ID): locally generated random identifier that remains stable across Research Runs in one Praxist environment; - `telemetry_run_id`: separate random identifier for one research run; - `event_id`: random identifier for correlation and deduplication of one event; - `event_sequence`: sequence number within the same research run; - `event_type`: one of the lifecycle events below; - `occurred_at`: client-side event time in UTC to second precision; - `error_summaries`: bounded structured error-category counts; and - `error_summaries_truncated`: whether the bounded error list was truncated. The random identifiers are not derived from usernames, accounts, device serial numbers, MAC addresses, IP addresses, hostnames, project paths, or task content, and are not simple incrementing identifiers. ### 2.2 Lifecycle events Praxist may emit four event types: 1. **`run_started`** records the generation ordinal, planned Peer count, and aggregate counts of Peers in planning, running, completed, cancelled, failed, and unknown states at the run-start boundary. 2. **`generation_finished`** records the same generation and aggregate Peer state fields at a durable Generation boundary. Peer state counts must sum to the planned Peer count. `completed` means only that the Peer lifecycle returned normally; it does not assert that a scientific result is valid. `unknown` means that a canonical terminal state could not be confirmed. 3. **`run_finished`** records active run duration in complete minutes and whether it reached the 43,200-minute (30-day) recording cap. Duration may be null when it cannot be determined. 4. **`run_reconciled`** records the same bounded duration fields when a previously unfinished run is resumed and trustworthy terminal processing is later completed. Individual Peer identifiers, outputs, prompts, research conclusions, and raw error details are not collected. ### 2.3 Structured error summaries Each lifecycle event may include up to 16 grouped summaries. Each group may contain only: - `scope`: `run`, `generation`, or `peer`; - `stage`: `setup`, `launch`, `execution`, `finalization`, or `reconciliation`; - `error_type`: `configuration`, `resource`, `orchestration`, `runtime`, `external_dependency`, `storage`, or `unknown`; - `error_code`: `PRX-CAPACITY`, `PRX-PEER-LAUNCH`, `PRX-PEER-RUNTIME`, `PRX-RUNTIME`, `PRX-RUN-FAILED`, or `PRX-UNKNOWN`; - `reason_code`: `auth_error`, `quota_exhausted`, `rate_limited`, `timeout`, `provider_unavailable`, `runtime_error`, `tool_unavailable`, `invalid_request`, `budget_denied`, `budget_expired`, `capacity_unavailable`, `process_start_failed`, `state_unreadable`, `unexpected_termination`, or `unknown`; - `count`: count of matching errors, capped at 65,535; and - `count_capped`: whether the count reached that cap. The structure cannot contain raw error messages, logs, stack traces, provider responses, or arbitrary text. ### 2.4 Time fields `occurred_at` is generated by the local Praxist client at a lifecycle milestone, converted to UTC using the local system clock, and may reflect clock inaccuracy. The Collector adds `received_at` after validating an event. This server receipt time cannot be supplied or changed by the client. It represents arrival time, not task completion time, and is used for storage management and retention rather than local duration calculation. ## 3. Data Praxist Does Not Collect Product-usage events do not include: - research task content, prompts, research results, files, filenames, project paths, or commands; - environment variables, API keys, saved login credentials, logs, stack traces, raw error messages, or arbitrary text; - model names, service-provider names, provider responses, or account information; - names, email addresses, operating-system details, hardware information, Python version, client time zone, cookies, or arbitrary request headers; or - individual Peer identities or individual Peer outputs. ## 4. Network Information Product-usage event bodies do not include IP addresses, cookies, or arbitrary request headers. Network services necessarily process connection information temporarily to deliver a request, protect the service, and apply rate limits. The Praxist Collector deployment disables access logging, strips forwarded IP, cookie, and client User-Agent headers before application processing, and does not persist them as product-usage event data. The client sends only a fixed, protocol-level User-Agent header. ## 5. Retention and Withdrawal Delivered raw events are retained for no more than 180 days and are then deleted by the scheduled retention process. The User may run: ```bash praxist product-usage withdraw ``` Withdrawal immediately disables future capture for the current user and deletes unsent local events. It does not delete already delivered events; those remain only until the scheduled retention period expires. The current state is available through: ```bash praxist product-usage status --json ``` During Agent-assisted OOBE, the supported explicit sharing replies are `Yes` and `Agree`; the supported refusal replies are `No` and `Disagree`. An Agent must not infer a choice from other language. --- # Documentation Policy The versioned Markdown under `docs/` is the only authored product documentation source. The static website, generated reference pages, search index, `llms.txt`, and `llms-full.txt` are derived from it. Praxist does not use a separately edited GitHub Wiki. ## One Fact, One Owner | Information | Sole owner | Other pages may | |---|---|---| | CLI arguments and defaults | `praxist.cli` parser code | Link to generated CLI reference | | Skill name and activation description | Each `skills/*/SKILL.md` front matter | Link to generated Skills reference | | Package installation, install commands, and filesystem effects | [Installation](https://praxist.sapient.inc/en/docs/getting-started/installation) | Link; the root landing page may show one canonical command | | User-facing first-run/OOBE sequence | [Quickstart](https://praxist.sapient.inc/en/docs/getting-started/quickstart) | Name the lane and link | | Agent-managed OOBE implementation | [Agent OOBE Runbook](https://praxist.sapient.inc/en/docs/agents/oobe-install) | Link without duplicating agent instructions | | Project prerequisites, research brief, and takeover stages | [Your First Task](https://praxist.sapient.inc/en/docs/getting-started/first-task) | Show one takeover invocation and link | | Goal-to-skill map and skill installation locations | [Agent Skills](https://praxist.sapient.inc/en/docs/user-guide/skills) | Name a relevant skill and link | | Task schema, precedence, and scientific ownership | [Task Projects](https://praxist.sapient.inc/en/docs/guides/task-projects) | Explain how a mechanism consumes the task contract | | Template/example distinction and example materialization commands | [Examples And Templates](https://praxist.sapient.inc/en/docs/guides/examples-and-templates) | Identify an asset and link without redefining the boundary | | Rocket Booster Recovery scientific details | `examples/rocket_booster_recovery/README.md` | Explain discovery and launch without duplicating its protocol | | Rocket Booster Recovery (Rust) scientific details | `examples/rocket_booster_recovery_rust/README.md` | Explain discovery and launch without duplicating its protocol | | Lifecycle semantics | [Direct CLI Operations](https://praxist.sapient.inc/en/docs/guides/operators) | Show one quickstart command and link | | Core/plugin/task boundary and artifact roles | [Architecture](https://praxist.sapient.inc/en/docs/concepts/architecture) | Summarize purpose and link | | Configuration ingress and precedence | [Configuration Discipline](https://praxist.sapient.inc/en/docs/concepts/config_discipline) | State a local input and link | | Agent-session and prompt-layout mental model | [Runtime Model](https://praxist.sapient.inc/en/docs/concepts/runtime-model) | Explain a mechanism-specific effect and link | | Runtime adapter capabilities | [Agent Runtimes](https://praxist.sapient.inc/en/docs/guides/agent-runtimes) | Identify a selected runtime and link | | API provider shapes | [API Providers](https://praxist.sapient.inc/en/docs/guides/model-providers) | Identify a selected provider and link | | Open-source model API shortlist and selection criteria | [Open-Source Model APIs](https://praxist.sapient.inc/en/docs/guides/open-source-model-apis) | Link without duplicating the shortlist | | Authentication and credential precedence | [Credentials](https://praxist.sapient.inc/en/docs/guides/credentials) | State a prerequisite and link | | Research-loop sequence | [Research Loop](https://praxist.sapient.inc/en/docs/guides/research-loop-variant-generation-flow) | Refer to a stage without redefining the sequence | | Maturity, close, incubator, peer-mix, and launch-freeze behavior | [Flexibility Controls](https://praxist.sapient.inc/en/docs/guides/research-loop-flexibility-controls) | Show task configuration only where Task Projects owns the combined profile | | Deep Innovation Gate (DIG) behavior and artifacts | [Deep Innovation Gate](https://praxist.sapient.inc/en/docs/guides/deep-innovation-gate) | State whether DIG is active and link | | Quality-Diversity (QD) allocation | [Quality-Diversity Allocation](https://praxist.sapient.inc/en/docs/guides/qdig-cohort-allocator) | Refer to the selected path and link | | Peer-session memory | [Peer Memory](https://praxist.sapient.inc/en/docs/guides/peer-local-structured-memory-long-context) | Refer to memory as context and link | | Experiment admission and resource ownership | [Central Experiment Scheduler](https://praxist.sapient.inc/en/docs/guides/central-resource-scheduler) | State the selected profile and link | | Tool catalog and frontier-tool behavior | [Tool Servers](https://praxist.sapient.inc/en/docs/guides/tool-servers) | Name a tool and link | | Literature source/provenance policy | [Scientific Literature Lookup](https://praxist.sapient.inc/en/docs/guides/scientific-literature-lookup) | State whether lookup is enabled and link | | Human-readable report triggers and semantics | [Run Reports](https://praxist.sapient.inc/en/docs/guides/user-facing-reports-and-init) | Link to a generated report or this guide | | Usage measurement and formulas | [Cost Estimation](https://praxist.sapient.inc/en/docs/guides/costs) | Link to measured artifacts and this guide | | Lossless token-saving mechanisms | [Cost Optimization](https://praxist.sapient.inc/en/docs/guides/cost-optimization) | State that a route uses the policy and link | | Contributor contract | `AGENTS.md` | Link without redefining it | | Software license | Root [`LICENSE.md`](https://github.com/sapientinc/praxist/blob/main/LICENSE.md) | Link without restating license terms | | User Agreement | [Praxist User Agreement](https://praxist.sapient.inc/en/docs/legal/user-agreement) | Link without restating legal terms | | Product-usage privacy policy, processing purposes, retention, and user rights | [Privacy Notice](https://praxist.sapient.inc/en/docs/legal/PRIVACY) | Link without restating the policy | | Exact versioned product-usage consent text | [Praxist User Data Collection Notice](https://praxist.sapient.inc/en/docs/legal/product-usage-data-notice) | The CLI loads this same package resource; other pages identify it and link | | Product-usage consent commands and collector operation | [Product Usage Controls](https://praxist.sapient.inc/en/docs/operations/product-usage) | Link to the operational procedure | | Product-usage implementation, endpoint, storage, and audit map | [Product Usage Technical Documentation](https://praxist.sapient.inc/en/docs/operations/DOCUMENTATION) | Link without duplicating implementation details | | Machine product-usage event schema | `praxist/product_usage/protocol.py` and checked-in JSON Schema | Technical and legal pages explain the contract; code remains authoritative for exact validation | | Hosted documentation URL | `praxist.cli.docs.DOCUMENTATION_URL` | Mirror it in checked package/site metadata and link to it | Tutorials sequence actions. Guides explain procedures. Concept pages explain mental models. Reference pages enumerate machine contracts. The root README is a product landing page, not a second manual. A minimal command may appear in a tutorial that needs the action, but its arguments, defaults, and edge cases remain in the generated reference. A page may summarize an adjacent contract only far enough to explain its own behavior; the summary must link to the owner instead of restating the contract. ## Generated Sources `scripts/build_docs_site.py` creates: - `docs/reference/cli.md` from the live CLI parser; - `docs/reference/skills.md` from skill front matter; - `site/llms.txt` as a compact machine-readable map; - `site/llms-full.txt` as the complete navigation-ordered corpus. Generated HTML and LLM exports are not committed. Generated Markdown reference pages are committed so package users can read them without building the site, but CI verifies that they exactly match their code-owned inputs. ## Navigation Ownership Every authored Markdown page must appear exactly once in `mkdocs.yml`. This makes its primary audience and information role explicit. The docs build rejects missing pages, duplicate navigation ownership, and broken local links. ## Build ```bash uv sync --extra docs uv run python scripts/build_docs_site.py uv run python scripts/build_docs_site.py --check-generated ``` The build runs in strict mode and does not contact API providers, read API keys, or start research services. ## Hosted Documentation Documentation validation runs on every pull request and every push to `main`. Successful pushes to `main` publish the generated site to the repository's private GitHub Pages project. Access follows repository read permission and requires GitHub authentication; the generated HTML remains derived output and is not committed. The repository variable `PRAXIST_PAGES_ENABLED=true` is the deployment switch. Maintainers can also trigger the `docs` workflow manually. Pull requests build and validate the complete site but never publish it. The canonical site URL is surfaced through `praxist docs`. Contributors use the local build commands above only to preview unmerged changes. --- # Legacy Migration Guide Legacy migration is allowed only when it preserves behavior while moving code toward the current core-plugin-task boundaries. ## Migration Pattern 1. Characterize the old behavior with tests. 2. Add the new protocol, plugin, or task boundary. 3. Route execution through the new boundary. 4. Add old-vs-new parity checks. 5. Preserve partial outputs and weak provenance when needed. 6. Delete the old path after the new path is proven. 7. Record migration context in the commit or pull-request description when the migration changes an architecture contract. ## What Not To Preserve Do not preserve obsolete package names, duplicate task catalogs, shell-owned semantics, fake production plugins, or hidden SAM-specific global defaults. Do not keep compatibility shims after they stop serving an active migration. ## Migration Tests Use characterization tests for old behavior, unit tests for the new interface, workflow smoke tests for run shape, and replay tests for artifact consistency. Long real GPU/API dogfood runs are manual/on-demand gates, not default unit tests.