Skip to content

About

Test local AI models like a scientist: ground-truth benchmarks for LM Studio + cloud LLMs. Vision/tools/reasoning/security, N=3 trials, SHA3-sealed evidence, live Rust+SSE dashboard. No marketing numbers — your hardware, your proof.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

962 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Calibration Scope

License: AGPL-3.0

Calibration Scope measures the gap between what a system states and what is actually true — machine-checked, sealed, and verifiable by anyone. Any subject. Any substrate. Silicon or carbon.

Every benchmark asks "what can it do?" Calibration Scope asks a sharper question: does what it says match what is actually so? A model states a verdict — we check it against a machine-verified logical ground truth. A run claims a result — we seal it with a SHA-3 hash you can re-verify. A system reasons in one voice or another — we measure whether the carrier changed the signal. The instrument never accuses; it measures a gap, continuously, and shows its work. Point it at a local model, a cloud endpoint, or a human taking the same battery — the method is the constant; only the subject changes.

If you study human reasoning: the same sealed battery measures carbon and silicon on identical stimulus. See the cognitive-construct crosswalk (DECISIONS.md §10.13, "Cognitive Atlas crosswalk").

Runs against local models (LM Studio) and cloud endpoints (Nous, OpenRouter, OpenAI, Gemini) — vision, tool use, reasoning, prompt-injection resistance — using ground-truth tests, N=3 trials, SHA-3-sealed evidence, and zero trust in anyone's marketing numbers. Built in Rust (Axum + Tokio + SQLx + PostgreSQL) with a single-file live dashboard driven by Server-Sent Events.

The defining feature: the same battery, the same seals, run against local and cloud models — then matched back and forth. Run a 30B model on your own GPU, run the same sealed battery on a cloud endpoint, and cross-reference the verdicts on identical stimulus. The methodology is the constant; only the silicon changes. That is how you answer "what am I dealing with today — and is it better than last week, or just bigger?"

And the loop closes on the human. The same fallacy taxonomy, the same N=3 discipline, the same sealed evidence can be pointed at you — take a calibrated logic/reasoning test, track your own weaknesses over time, and watch your profile move. Silicon and carbon under one instrument. That is the point: a tool to honestly measure intelligence wherever it shows up.

Made by IT Help San Diego Inc. · A project of the Intellectual Resistance

Status

  • In active development — scientific validation is the main bottleneck, not features.
  • Hermes-ready: integrated in Hermes Desktop as of July 2026. Model routing, cloud key setup, and local clean-room execution are verified live.
  • Standalone-capable: runs on any macOS/Linux box with Rust + PostgreSQL + LM Studio. No Hermes dependency.
  • Modular architecture: backend exposes REST + SSE; frontend is a single static HTML file. MCP server layer is planned so external tools (OpenClaw, bots, scripts) can drive benchmarks programmatically.
  • Public beta: the core pipeline works (clean-room, blind tests, SHA-3 seals, speculative decoding measurement). The science — competitive cross-reference, fallacy taxonomy expansion, contamination resistance — is being validated now.

📊 Published Findings

docs/FINDINGS.md — the living record of what we've measured, sealed with SHA-3 provenance. Current results:

  • Carrier Color: a model's verdict tracks the carrier of identical logical content (English prose vs Lean formula vs haiku vs flattery), not the signal. Heavy carriers drag a strong model: Lean formulas and flattery both cost ~8 points against an unscaffolded baseline (99.0% → 91.2%, Fisher p = 0.019, N = 102/arm) — both inverse hypotheses ("Lean is clean", "flattery lifts") falsified. The relative ordering among carriers is not resolvable at this N (adjacent steps p = 0.50–1.00); the pre-registered paired design (docs/experiments/carrier_color_experiment_spec_v1.md) is built to settle it.
  • Carrier-immunity threshold (candidate): the small e2b shows a real carrier drop (99% → 91% under Lean/bribe — the endpoint separation is statistically supported). Larger models (nemotron 30B, Fable 5) show no carrier sensitivity we can resolve at current N (all at/near the 100% ceiling — small-N ceiling values, not a proven 100%). Whether immunity reflects a real capability/headroom threshold, independent of substrate (local vs cloud), is exactly what the pre-registered paired-design experiment (N≈420, docs/experiments/carrier_color_experiment_spec_v1.md) is built to answer — not yet a settled result. We hold our own claims to the standard we measure others by.
  • Verified leaderboard: local models ranked on the reasoning battery, clean post-fix runs. Goldilocks floor: <1.5B breaks, 1.5B barely, 2B (gemma-4-e2b) genuinely usable.
  • The "free bot" honesty check (Fountain/Trickle/Mirage): "free" is not one thing. FOUNTAIN (flows freely) vs THROTTLED (free but rate-limited) vs MIRAGE (claims free but fails). Look in the gift horse's mouth.

Every number is a real, sealed run — nothing derived, estimated, or marketing.


Why this exists

Standardized benchmarks (MMLU and friends) are public — models have trained on them, so the numbers can't be trusted alone. Meanwhile, people running local AI on their own machines have no honest way to answer basic questions:

  • Which of my models can actually read text in a screenshot — and which one fabricates plausible-sounding text when it can't?
  • Which model should sit in each job slot (vision, tool routing, command approval) based on my measured evidence, not parameter-count intuition?
  • Is a "failure" a real capability gap, or my infrastructure lying to me?
  • And the same questions, pointed at me: where are my own logical blind spots, and are they moving?

The local/cloud split is the feature, not a limitation. Run the identical battery against a model on your GPU and against a cloud endpoint; the cross-reference tells you whether the cloud model is genuinely better or just more expensive. And the human-calibration path turns the same instrument on its operator — because the goal was never "rank the bots." It was: honestly measure intelligence, silicon or carbon, wherever it shows up.

This tool answers those questions with the discipline of a lab notebook:

Principle Implementation
No answer leakage Ground truth lives only in the database; the model is never shown what it's scored against. The test builder rejects prompts that contain their own answer.
N=3 trials, always One pass can be luck. Capability axes score PASS / FAIL / INTERMITTENT; the security axis scores RESISTED / COMPLIED / INTERMITTENT. All three carry UNTESTED when nothing was measured — absence of evidence is not a verdict. Note the security words are past-tense and about the run, not adjectives about the model: a model that RESISTED is not thereby safe, and one that COMPLIED may simply have been misread by the grader (this has happened — see src/executor/scoring.rs).
Objective scoring Verdicts come from comparing output to ground truth (exact match, substring, regex, spatial relations, tool-call shape) — never from asking a model its opinion.
Clean-room runs Before each local run, every loaded model is ejected, only the target is loaded, and RAM residency is verified by polling — never assumed.
Sealed evidence Every trial stores the exact prompt, raw response, reasoning trace, and latency. Every run is sealed with a SHA3-512 provenance hash. Test images are SHA3-256-pinned so the stimulus can't drift.
Infra errors ≠ capability failures A config bug that blocks requests is recorded as infrastructure noise, excluded from the capability score — a model is never blamed for your network.
Full citation graph Every trial row links to its exact test (prompt + pinned image + ground truth) and its run seal. Query the evidence in both directions.

What it caught in its first week (real examples)

  • A 12B vision model confidently fabricated an entire sidebar of a screenshot it couldn't read — invented plausible note titles, zero of which existed. The token receipt (whole image compressed to ~256 vision tokens) explained why: the text was physically unreadable at that budget, and the model invented rather than admitting it.
  • A 2B model failed the same test deterministically — same wrong answer, three trials in a row (it read a menu-bar icon and reported the wrong app). Four other models read the same pixels correctly, 3/3 each. That one screenshot is now a permanent regression test.
  • The leaderboard formula itself was caught ranking a text-only coder model #1 overall despite a 100% hard-fail on vision — because the old score only counted wins. Fixed, regression-tested, documented in the commit history.
  • An empty model response (reasoning budget exhausted, finish_reason: length) was being rendered as if it were an answer. Now it's a loud failure banner: "NO FINAL ANSWER — this is a failure, not a result."

The commit history is deliberately forensic — most fixes cite the live incident that motivated them.

Features

  • 🏁 Benchmark grid — every model in your LM Studio library plus configured cloud models (Nous, OpenRouter, OpenAI, Gemini), per-axis verdicts with latency, live SSE telemetry (ejecting → loading → resident → trial → verdict), real timestamps, no spinners anywhere.
  • 🔁 Local ↔ cloud matching — the same battery and the same SHA3 seals run against local and cloud models, so you can cross-reference verdicts on identical stimulus and see whether a cloud model is genuinely better or just more expensive.
  • 🧠 Human calibration (in progress) — the same fallacy taxonomy, N=3 discipline, and sealed evidence aimed at you: take a calibrated logic/reasoning test, track your own weaknesses over time, and watch your profile move. Silicon and carbon under one instrument.
  • 🏆 Loot page — leaderboard + "recommended squad" (best verified model per job slot), and a capability router that assigns primary/fallback models per axis from lifetime evidence, with stated reasons and evidence links (/api/router/plan).
  • 🧪 Prompt Builder — side-by-side workbench: compose (text + image) on the left, results on the right. Reasoning traces shown separately from committed answers. Persistent run history — every test you run is kept, queryable, revisitable. Prompt-length checker with instant heuristic + optional exact token count.
  • 📋 Test Registry — blind by default (ground truth requires an explicit audit view), custom test builder with anti-leakage validation, viewable SHA3-pinned image attachments.
  • 🖥️ Reality check — the setup page measures your machine (RAM via sysctl, GPU ceiling via Metal's recommendedMaxWorkingSetSize — a documented API, not folklore, live memory pressure, LM Studio state) and computes an honest AI RAM budget with the formula shown. Every number carries the command it came from.
  • ⚙️ Hermes-aware — if you run Hermes Agent, the setup page reads your actual config (allowlisted fields only — never credentials) and shows verified ✅ state for main model and auxiliary task slots.
  • 🤖 MCP server — a real Model Context Protocol server at POST /mcp (JSON-RPC 2.0). A bot can connect, discover 11 tools (tools/list with JSON-Schema args), and tell Calibration Scope to do stuff: run_benchmark (returns run_id immediately), get_run (poll state), abort_run, list_models (with verdicts + size_gb), get_model_verdict, get_leaderboard, get_carrier_color, get_owl_state, get_test_spec, list_tests, get_status. Every tool is documented + verifiable (learning from LM Studio's API anti-patterns — no hidden state, no "maybe it works" endpoints, honest data). See docs/mcp-server-design.md.

Using this with Hermes Agent / Hermes Desktop

This dashboard pairs naturally with Hermes Agent (Nous Research's open agent runtime) — it was built alongside a live Hermes deployment:

  • Route by evidence, not vibes: run the benchmark battery, then open GET /api/router/plan — it assigns a verified primary + fallbacks per capability axis. Map those onto Hermes' Settings → Model Settings → Auxiliary Tasks slots (vision, MCP tool routing, approval classification, web extract). The Setup tab shows the exact mapping.
  • The "Approval" slot is your security surface: Hermes' smart auto-approve sends shell commands to a model for APPROVE/DENY/ESCALATE judgment. This dashboard's security axis measures exactly that job — prompt-injection resistance with real injection payloads — so you can pin a proven local model there instead of guessing.
  • Config verification, read-only: the Setup tab reads ~/.hermes/config.yaml through a strict allowlist (never credentials) and shows live ✅/⚠️ against what Hermes is actually configured to do.

Hermes is not required — the dashboard works standalone with LM Studio and/or any OpenAI-compatible cloud endpoint.

Stack

  • Backend: Rust — Axum 0.8, Tokio, SQLx 0.9, PostgreSQL, reqwest, SHA-3. One static binary.
  • Frontend: one HTML file, zero frameworks, zero build step. SSE for live updates.
  • Evidence store: PostgreSQL (inspectable with any SQL client — the schema is the API).
  • Model I/O: LM Studio REST (/api/v0), OpenAI-compatible cloud endpoints.

No telemetry. No external calls except the model endpoints you configure. Binds to 127.0.0.1 only.

Install

macOS, one line — installs the self-contained binary and PostgreSQL, sets up a launchd service that survives reboots, and opens the dashboard:

curl -fsSL https://calibrationscope.com/install.sh | sh

Homebrew directly — same payload (the formula brings postgresql@17 with it); you wire up DATABASE_URL yourself:

brew install it-help-san-diego/tap/calibration-scope

Linux — prebuilt binary via the release installer (tag-pinned URLs in docs/RELEASING.md); PostgreSQL from your distro.

Prebuilt binaries are self-contained — the dashboard and the SHA3-pinned test stimuli are embedded (src/embedded.rs) — and every release carries GitHub build-provenance attestations (gh attestation verify). Local-model runs need LM Studio serving on :1234; cloud runs just need keys.

Hardware reality, disclosed: the instrument is featherweight and runs on nearly anything, ARM included — hosting models is the heavy part, and the iron deserves honest labels: small machines top out around the 2B floor our own leaderboard measured (FINDINGS); "local AI at frontier quality" means something like a maxed-out MacBook Pro — a several-thousand-dollar supercomputer in a laptop costume; and "cloud" means renting someone else's data center by the token. On weak hardware you're benchmarking cloud models — and yourself. The scope runs on a $60 board; the minds it measures do not.

Updating

No auto-updater, by design — the instrument never phones home, and that includes not checking for its own updates. Updating is one deliberate line:

brew upgrade it-help-san-diego/tap/calibration-scope && launchctl kickstart -k gui/$(id -u)/com.calibrationscope.dashboard

(The kickstart matters: brew swaps the binary on disk, but the running service keeps the old one in memory until restarted.) If you installed via the release tarball instead of brew, re-run the newer release's installer. Every release is announced on the GitHub releases page — watch the repo to get notified.

Quick start (from source)

# Prereqs: Rust toolchain, PostgreSQL, LM Studio with its local server on :1234
git clone https://github.com/IT-Help-San-Diego/calibration-scope.git
cd calibration-scope
cp .env.example .env          # set DATABASE_URL (and optional cloud API keys)
cargo run --release           # migrations run automatically — .env loads from the working directory
# open http://127.0.0.1:8768

Sync your LM Studio library from the dashboard (LM Studio → Sync), pick a model, click ▶ Run — and watch the live log. Verdicts land on the grid with latency; evidence lands in Postgres with a seal.

Operations

This repo includes a launchd-managed backend on macOS. Use these commands instead of ad-hoc cargo run processes.

Service

Name: ai.hermes.calibration-scope-dashboard Binary: ~/Documents/GitHub/calibration-scope/target/release/calibration-scope-dashboard Port: 8768 on 127.0.0.1 Database: postgres://<dbuser>:<dbpass>@localhost:5432/calibration_scope Logs: /tmp/calibration-scope-dashboard.out, /tmp/calibration-scope-dashboard.err

Start / stop / restart

launchctl start ai.hermes.calibration-scope-dashboard
launchctl stop ai.hermes.calibration-scope-dashboard
launchctl kickstart -k gui/$(id -u)/ai.hermes.calibration-scope-dashboard

Health

curl http://127.0.0.1:8768/api/status

Run control from the backend API

# start a run
curl -X POST http://127.0.0.1:8768/api/runs \
  -H 'content-type: application/json' \
  -d '{"model_key":"google/gemma-4-31b-qat","axes":["vision","reasoning"],"load_mode":"clean-room"}'

# list runs
curl http://127.0.0.1:8768/api/runs

# run detail
curl http://127.0.0.1:8768/api/runs/628

# abort an in-flight run
curl -X POST http://127.0.0.1:8768/api/runs/628/abort

Database access

Preferred: TablePlus connection 127.0.0.1:5432 → database calibration_scope → your configured PostgreSQL user. The schema is the API: every trial row links to its exact test and run seal.

Troubleshooting

  • If the binary was rebuilt, use launchctl kickstart -k instead of launchctl start so launchd loads the new executable.
  • If runs show status = error but trials exist, the executor preserved partial evidence; the database still contains complete trial results for post-mortem analysis.
  • Quarantined runs are excluded from leaderboard/router scoring by default; review them via /api/quarantine when needed.

Philosophy

This project believes the flood of AI-generated junk science gets fixed by making rigorous method cheap, not by gatekeeping. Everyone with a laptop and curiosity can run a controlled experiment: pinned stimulus, committed answers, N=3, sealed results. The dashboard is deliberately a teacher — it explains its formulas, shows its receipts, and marks its heuristics as heuristics.

If a number on the screen can't cite where it came from, that's a bug. File it.

License & attribution

Calibration Scope is licensed under the GNU Affero General Public License, Version 3.0.

This is a fork point, not a history change: everything published through v0.1.0-beta.1 shipped under Apache-2.0, and that grant is irrevocable — those versions remain Apache-2.0 in perpetuity for anyone who has them. AGPL-3.0 applies from the next release forward.

  • Copyright © 2026 IT Help San Diego Inc. — sole owner (contributors: Carey James Balboa; tool-authored commits carry no independent copyright standing).
  • Research published under Carey James Balboa and IT Help San Diego Inc., as part of the Intellectual Resistance program.
  • The benchmark methodology, test battery, scoring logic, and SHA3-provenance design are original works of IT Help San Diego Inc.
  • AGPL-3.0 is OSI-approved and machine-classified: dependency scanners and academic license whitelists accept it cleanly. The one obligation it adds over Apache-2.0 is the anti-shelving tooth — operating a hosted modified derivative requires publishing your changes, so a fork can never be withdrawn into private.
  • Trademark: "Calibration Scope" and "IT Help San Diego" are trademarks of IT Help San Diego Inc. The Owl of Athena is a historical/public symbol used for thematic identity and is not claimed as a trademark. The license does not grant rights to use the Calibration Scope or IT Help San Diego marks except for reasonable attribution.
  • See NOTICE for attribution and trademark details.

About

Test local AI models like a scientist: ground-truth benchmarks for LM Studio + cloud LLMs. Vision/tools/reasoning/security, N=3 trials, SHA3-sealed evidence, live Rust+SSE dashboard. No marketing numbers — your hardware, your proof.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages