Star 历史趋势
数据来源: GitHub API · 生成自 Stargazers.cn
README.md

DeepGraph: an autonomous research agent that improves by being judged by reality

The autonomous research agent of JouleBeat, an open-source RSI lab.
It reads the literature, finds open questions, runs the experiments,
and lets reality decide what is true.

Website · Results · Showcase · Architecture · 中文

MIT Python 3.12+ Harness × Benchmark


Why DeepGraph

JouleBeat builds self-improving AI by letting harnesses and benchmarks evolve together.

How an agent works (its harness: tools, memory, workflow, verification) now moves capability by as much as a model generation, and open source has made harness variants cheap to generate. What is still scarce is the benchmark that tells real progress from a lucky score. So we put both in the loop: harnesses compete under benchmarks, and benchmarks compete on how well they predict results that arrive later.

That loop needs a steady stream of real tasks whose answers can be checked. Computational research is exactly that: a conclusion is confirmed or refuted by rerunnable code, held-out data or an independent recomputation, in minutes to days. DeepGraph is where those tasks come from.

  • For researchers, it is a research agent that delivers conclusions you can check.
  • For us, it is the accelerator of our own R&D: every task is a judged trial for the next harness and the next benchmark.

How it works

StepWhat happens
1ReadHarvests arXiv at scale and extracts claims, methods and results
2MapMerges them into an evidence graph of entities, relations and contradictions
3AskTen structural detectors find open questions in pure SQL, with no LLM in the path
4TestTurns a question into a pre-registered, budgeted experiment on CPU or GPU
5JudgeAn evidence ladder issues a verdict only when the comparison is fair, and the result updates what gets searched next

Every verdict is auditable, and the system audits itself: in August 2026 it found and retracted 11 of its own conclusions that rested on empty model outputs. A search loop is only as good as the filter that decides what is true. Details in Results.

At a glance

Live production data, August 2026.

Papers24,407 harvested, 7,005 through the full evidence pipeline
Evidence graph248,441 entities, 735,913 relations, 238 contradiction clusters
Discovery25,652 mapped research opportunities
Experiments150 audited verdicts across six research agendas
Compute921M LLM tokens invested

The JouleBeat stack

LayerWhat it isStatus
DeepGraphResearch agent: real tasks in, checkable conclusions outIn service, with paying research groups
Harness × Benchmark librariesVersioned harnesses, and benchmarks scored on how well they predicted later outcomesHarness versioning running; benchmark validation in build
Evolution KernelThe engine: propose, evaluate, accept or roll backOpen source, pip install evolution-kernel

Quick start

Python 3.12+ and an LLM API key are the minimum; PostgreSQL enables the full control plane.

python3.12 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env    # set at least DEEPGRAPH_LLM_API_KEY
export $(grep -v '^#' .env | xargs)
python3.12 main.py      # open http://localhost:8080

Full deployment (PostgreSQL, systemd, remote GPU backends): docs/DEPLOY.md.

Learn more

  • Architecture: design principles, the loop, evidence gates, repository map, tests
  • Results: every measured outcome, read directly from the production database
  • Showcase: demo route and case studies
  • Roadmap

Contributing

Researchers: bring a computational question at deepgraph.joulebeat.com. Developers: issues and pull requests are welcome; contributions are covered by the CLA.

License

MIT

关于 About

Token-scale scientific discovery engine — autonomous hypothesis generation, experiment execution, and knowledge graph synthesis

语言 Languages

Python92.2%
JavaScript4.0%
CSS1.9%
HTML0.7%
Shell0.5%
TeX0.5%
TypeScript0.2%
PowerShell0.0%

提交活跃度 Commit Activity

代码提交热力图
过去 52 周的开发活跃度
522
Total Commits
峰值: 170次/周
Less
More

核心贡献者 Contributors