Two-Tier AI Compute

Reasoning complexity determines the model

Model Routing Architecture

Expensive Opus tokens are spent only where adversarial review and final decisions are made.

FRONTIER — Claude Opus

Deep adversarial reasoning and final approval. The two most important decisions in the pipeline: challenging the trade (Deep Thinking) and signing the order (Manager Agent).

OPUS $5 / $25 per 1M tokens 2 nodes

SYNTHESIS — Claude Sonnet

All analysis, synthesis, planning, and control work. The muscle of the pipeline — carries 91% of tokens at 60% of Opus's unit price. Handles data integration, market analysis, research, fact-checking, trading plans, backtesting, position sizing, risk management, and monitoring.

SONNET $3 / $15 per 1M tokens 12 nodes

NON-LLM — Deterministic Compute

Pure technical ingest and deterministic execution. Data Sources (WebSocket/REST feeds) and the Execution Agent (order placement) never call an LLM and consume zero tokens. A hard design rule: AI thinks, code executes.

NON-LLM $0 2 nodes

Token Distribution Per Run

One full-market run: ~6.94M tokens, $28.48

SONNET — 12 AGENTS

91% of Tokens · 85% of Cost

6,327,000 tokens per run. $24.11 per run. Handles the entire analysis pipeline from data integration through risk management and monitoring.

6.3M tokens $24.11 12 agents
OPUS — 2 AGENTS

9% of Tokens · 15% of Cost

615,000 tokens per run. $4.38 per run. Owns the two most critical decisions: adversarial challenge and final signed approval.

615K tokens $4.38 2 agents

Optimization Strategies

Maximizing intelligence per dollar spent

Two-Tier Routing

Sonnet carries 91% of tokens at 60% of Opus's unit price. Opus is reserved exclusively for the two hardest jobs: adversarial review and final approval.

Early Exit

93% of runs (NO SIGNAL) stop after the Researcher — $18.90 instead of $28.48, saving ~34% per run. The system actively says NO to conserve resources.

Conditional Opus Activation

Deep Thinking runs only 9 of 1,440 daily runs. All of Opus costs ~$43/day (0.2% of budget). NO SIGNAL and NO TRADE runs never reach it.

Math Outside the LLM

Backtest, position sizing, and execution run on numerical engines — the LLM only reads summarized results. Zero tokens for heavy computation.

Early Fact-Checking

Rejecting junk evidence at step 5 reduces input for every downstream agent. Every rejected item saves tokens at every subsequent step of the pipeline.

Prompt Caching

System prompts and schemas are fixed across runs — 50–70% of input tokens are cacheable, pushing real cost well below the nominal budget.

Data Sources

8 live feeds, 42 symbols, 1,240 events/second

Binance Perpetuals
OKX Perpetuals
Bybit Perpetuals
Glassnode On-Chain
X / Twitter Stream
Telegram KOL
CoinDesk
Bloomberg

$10M/year of AI compute. 2.74 trillion tokens/year.

A 13-expert trading desk powered by two-tier intelligence architecture.

Request Access