One task. A bounded scope.
Give an agent a research question, source material, and an explicit stopping condition. Dependencies determine which work can start.
Coordinate coding, evidence gathering, and independent review across long-running projects. SciDocket keeps tasks and research decisions connected to their source files.
For computational labs and R&D teams
Specify the source data, fixed assumptions, expected outputs, and checks required before the result can be accepted.
Select a stage to explore. An illustration of the prototype workflow, not a live run.
Research needs more than a generated answer. SciDocket connects each assignment to its evidence, code changes, checks, and review.
Give an agent a research question, source material, and an explicit stopping condition. Dependencies determine which work can start.
A reviewer checks the result against the task and project requirements. Requested changes return to the same recorded work.
The comparison is reproducible. Explain the treatment of missing observations before this result is adopted.
Persistent task state, retained workspaces, and bounded retries support recovery when a worker or session stops.
Researchers set project objectives and assumptions. Material changes outside the assigned scope require a human decision.
A persistent workflow connects the tools. GitHub keeps the task, source files, and review history together.
Record the question, constraints, dependencies, and evaluation criteria in a GitHub task.
Route eligible work to a configured coding or research backend. Local workers use separate workspaces.
Run the relevant tests and route work for independent review before adoption under project rules.
Preserve the changes, decisions, and limitations. Approved summaries can also be published to Notion.
Start with one bounded workflow in an existing research project. These examples describe intended uses, rather than customer results.
Assign a focused source review. Record definitions, inclusion decisions, provenance, and unresolved gaps in files that another researcher can inspect.
Research tasks often combine repository context, code changes, tests, and written interpretation. The Research OS prototype supports Claude Code as an execution and review backend within that process.
SciDocket is model-independent. Backend availability depends on configuration and access. Claude integration does not imply Anthropic sponsorship.
Bounded repository tasks, structured handoffs, and separate execution and review phases. The host checks returned changes before they enter the GitHub workflow.
Evaluate programmatic execution and review with explicit budgets and task-specific tests. Measure reliability and researcher effort before extending access to more teams.
SciDocket is being developed from Research OS, a private, GitHub-based research orchestrator. The current implementation includes a local runtime, worker interfaces, review handling, and operational views.
Issues, dependencies, branches, and pull requests retain task context and review history.
Task claims, retries, execution receipts, and restart recovery support long-running work.
The host manages Git changes and workflow transitions. Local workers return bounded outputs for validation.
Slack provides operator input; a read-only dashboard exposes status; an optional Notion integration publishes retained summaries.
SciDocket addresses a practical coordination problem: using AI across projects while retaining the assumptions, source material, and checks needed to evaluate the results.
The goal is to make this workflow usable by other computational research groups and small R&D teams.
FOUNDER / RESEARCHER
A researcher at the University of Cambridge, working on material supply chains and computational modelling. His background includes stock-and-flow models and evidence-intensive environmental research.
SciDocket is an independent project. The founder's university affiliation describes his background and does not imply institutional endorsement.
The software is a working private prototype. We are seeking conversations with computational labs and R&D teams about a bounded pilot. Team onboarding and broader access are still being developed; there is no public self-service signup.
Agents can propose methods, execute tasks, and challenge assumptions within their assignment. The project defines which results can be adopted. Material changes of scientific direction outside the authorized scope require a human decision. Review does not guarantee scientific correctness.
The prototype runs from a local host and records work in GitHub. AI providers may receive the content needed for assigned tasks, according to the configured backend. Local orchestration does not mean that all processing stays on your machine. Confidential data and institutional requirements must be assessed before a pilot.
No. The prototype supports configured coding backends such as Claude Code and Codex, alongside explicit handoffs to browser-based research workers. The proposed Claude API integration is future work. Models and accounts do not replace the task record or the review process.
A useful first pilot would test one research workflow: whether outputs can be reproduced, whether another researcher can inspect the evidence, how often intervention is needed, and how failures are recovered. These are evaluation objectives, not performance results already established.
Tell us what your team is working on, which tools you use, and where coordination becomes difficult.