AI Inference on PYRAX

AI inference is the runtime side of AI — serving trained models to answer prompts, score data, and drive agents — and it is now the dominant, recurring cost of production AI. PYRAX turns inference into a verifiable, pay-per-use market through PYRAX Compute: jobs are priced in compute units at a fixed 8 PYRX/CU, funded by on-chain escrow, and settled against a single ComputeReceipt whose result is checked by a 4-rung verification ladder rather than trusted blindly.

Market size
$76B (2024)
Projection
$255B by 2030 · ~22% CAGR

The market

$76B
AI inference market (2024)
~60–70%
Share of AI compute spend on inference
65%
Enterprises deploying GenAI in production
3–5×
Cost premium vs. commodity GPU rental

Source: Grand View Research, 2024. Figures are indicative and provided for context.

What's broken today

Inference bills scale with every request, so serving costs quickly dwarf one-time training costs.
You cannot prove a hosted API actually ran the model and settings you paid for — quantization, routing, and caching are invisible.
Sending prompts and documents to a third-party endpoint leaks proprietary and personal data.
Capacity is concentrated in a few clouds, creating price power, waitlists, and single points of failure.

How PYRAX transforms it

Concrete network elements mapped to this business.

PYRAX Compute verifiable inference (pay-per-CU)

Each request becomes a metered job priced at a fixed 8 PYRX/CU and bound to a ComputeReceipt that records the model, inputs digest, and output, so you pay only for the work actually performed.

4-rung verification ladder

Results climb from provider attestation, to redundant re-execution by independent workers, up to ZK proofs — you choose how strong a guarantee each job needs instead of trusting the endpoint blindly.

On-chain escrow settlement

Funds are locked in escrow at submission and released against a verified ComputeReceipt, so honest workers are paid atomically and disputed or failed jobs are refundable.

Compute-to-data + shielded transfers

Inference runs next to private data with prompts and outputs moved as shielded transfers, and a scoped viewing key lets an auditor confirm what ran without exposing the payload.

Local-first consumer-GPU workers

An RTX 3060-baseline runtime lets anyone contribute a GPU, and VRAM/trust-tier scheduling routes each model to a worker that can actually hold it, widening supply beyond the hyperscalers.

Buildathon: dApp ideas

Ship-ready concepts for ai inference on PYRAX.

ProofServe

01

Verifiable inference gateway that returns each completion alongside a ComputeReceipt and a caller-selected verification rung.

ComputeComputeReceipt

EscrowInfer

02

Pay-per-CU API where callers escrow PYRX per request and funds release only when the result verifies.

ComputeEscrow

PrivatePrompt

03

Shielded inference relay that hides prompts and outputs on-chain while issuing viewing keys to compliance reviewers.

ShieldedPYRAX Compute

ZKOracle

04

ZK-proven model oracle that feeds smart contracts scores it can prove were produced by a specified model.

ZKPYRAX Compute

RouterMesh

05

Trust-tier router that spreads inference across independent workers and re-executes a sample for redundant verification.

ComputeScheduler

AgentLedger

06

Autonomous-agent runtime whose every tool call and model step is metered and receipted for auditable spend.

ComputeEVM

EdgeServe

07

Consumer-GPU inference pool that lets home RTX operators earn PYRX for served CUs with VRAM-aware scheduling.

ComputeEdge

FairDecline

08

On-chain decision endpoint that attaches a verifiable receipt to model outputs so a disputed answer can be reproduced.

ComputeReceiptZK