Whitepaper · Draft 0.1

Proof of work that does the work.

Vearl secures a blockchain with the same GPU computation that answers a client's AI request. This paper describes the protocol, the consensus rules, the compute market, how results are verified, the cryptography, and what is still unsolved.

30 September 2026 Testnet specification Not audited Testnet tokens have no value
Contents

00 · AbstractA chain secured by useful computation.

Vearl is a proof-of-work blockchain in which the work that secures the chain can also be a paying client's computation. Miners multiply 8-bit integer matrices on GPUs. Every 16 × 16 tile of a noised product is a lottery ticket, and a ticket below the difficulty target is a block.

A block counts as useful only when the matrices are the ones a client committed to in an escrowed job. The client checks the returned result with a probabilistic test that costs a small fraction of the computation itself. Any single wrong value can be proven on chain with a proof of about a kilobyte, and the worker that returned it loses its bond. When there is no client demand, miners work on matrices derived from the protocol itself, so the security of the chain never depends on the market.

Results are exact integers, so every node computes identical bits and a fraud proof is never ambiguous. On this base, programs of up to 240 tensor operations run as verified jobs, and jobs chain together. The testnet has verified a 12-layer vision transformer, a 110-million-parameter text encoder and a small language model that writes stories word by word, each checked bit for bit against an independent reference engine.

30 s

target block time, adjusted by ASERT

2−24

worst-case error of the client's 24-round Freivalds check

240

operations in one verified tensor program

0

external audits so far

Status

Vearl is a closed testnet. The code has not been reviewed by a third party, and several security arguments in this paper are empirical or simulated rather than proven. Testnet tokens have no value, are not for sale and must not be treated as an investment. Nothing in this paper is financial advice.

Author. Clovis Anicet, founder of Vearl, a publicly identified (doxxed) member of the project. Contact: contact [at] vearl.org.

How to read this paper

Claims carry a label that says how strong the evidence is. The labels are used throughout.

Implementedin the reference node and covered by tests
Measuredrun on real hardware, numbers reported
Simulatedresult of a reproducible model, not a proof
Hypothesisassumed, not demonstrated
Opennot solved
Literatureestablished result from published work

Vearl is the name of the project and of its reference implementation. The testnet unit is divisible into 108 base units. A final token name has not been chosen.

01 · IntroductionWhy hashing is not the only way.

1.1The cost of hashing

Proof of work secures a ledger by making block production expensive. An attacker who wants to rewrite history has to out-spend everyone else, and that cost is what honest participants rely on. In Bitcoin the expense is SHA-256 hashing, and the result has no use outside the block.

Replacing hashing with computation that somebody actually wants is an old idea. It runs into two hard problems.

1.2Two problems with useful work

Verification must be cheap. Every node must be able to check that a block contains valid work, and a client must be able to check its own result, without redoing the computation. Matrix multiplication is one of the few heavy workloads where this is possible: a product of two n × n matrices costs about n3 operations, while a probabilistic check costs about n2 Literature [4].

Usefulness cannot be proven. A protocol can check that a product was computed correctly. It cannot check that anybody wanted it. If the miner chooses its own input matrices, a miner that feeds random numbers through the whole pipeline is accepted exactly like one that runs a model. A recent empirical study of a deployed network argues that this is what happens in practice [2]. We cite it as one third-party study whose numbers we have not reproduced.

1.3The Vearl approach

A protocol cannot prove that a computation is useful to a third party. It can prove that someone paid, and burned a share of the payment, for that exact computation to be done. Vearl calls this fee-bound useful work and makes it the only certificate of usefulness the protocol claims.

  • The input matrices of useful work are committed by a client in a job transaction that locks the client's payment in escrow. A miner can earn a useful-work bonus only for matrices a client committed to.
  • The bonus is capped by the amount the client burned. Creating fake jobs for yourself therefore never pays better than mining ordinary work.
  • The computation that answers the client is the same one that produces lottery tickets. A GPU running a client's job mines at essentially the same rate as a GPU running protocol-derived matrices, so client demand adds to security instead of competing with it. The measured overhead is about 9 %; the equal-rate assumption still has to be confirmed on a real network Hypothesis.
  • With no client demand, miners use protocol-derived matrices. Security does not depend on the market, and no usefulness is claimed for that work.

1.4Design goals

  1. Security with an empty market. Every valid block earns the same base reward whatever the work class.
  2. Cheap, exact verification. Nodes recompute one tile per block. Clients check a result in quadratic time. One wrong value is provable with a short proof.
  3. Deterministic integer semantics. No floating point anywhere in consensus. Two honest workers produce identical bits, which is what makes a fraud proof indisputable.
  4. Honest economics. Incentives are bounded by burned value, and every parameter is labelled provisional until it is calibrated.
  5. A small, auditable surface. Standard cryptographic primitives only, a byte-level specification, and independent second implementations that are checked against the node on every change.

1.5Related work

The noisy matrix-multiplication proof of work used here builds on the construction of Komargodski, Schen and Weinstein [1], in which low-rank noise derived from a public seed turns an arbitrary matrix product into a lottery while keeping the honest overhead close to zero. Vearl keeps that idea and changes who owns the inputs: they come from escrowed jobs, not from the miner. Pass [3] studies the economics of useful proof of work and identifies a regime in which block rewards subsidise the useful computation; Vearl's separation between security (block reward) and usefulness (client fees) follows that reading. The difficulty rule is ASERT [8], the optional reorganisation brake follows ECIP-1100 [9], and account signatures include the post-quantum scheme standardised in FIPS 205 [5].

02 · System overviewTwo kinds of work, one chain.

2.1Participants

Who does what
Client
Commits input and weights, locks an escrow, receives the result, checks it, and accepts it or proves it wrong.
Worker
Claims a job, posts a bond, computes the result, mines blocks with the client's matrices, and publishes commitments to its intermediate results.
Miner
Any node that produces blocks. A worker is a miner that also holds client data. A miner without a job uses protocol-derived matrices.
Validator
Every full node. Recomputes one tile per block, applies transactions, and enforces every rule in this paper.
Challenger
Anyone. Submits a fraud proof against a wrong result and earns half of the cheater's bond.
Light client
Checks block headers, the proof of work and account balances without the full state.

2.2Life of a job

One useful job, end to end
  1. ClientCommits matrices, locks escrow
  2. →
  3. WorkerClaims, posts a bond
  4. →
  5. WorkerMines with the client's matrices
  6. →
  7. WorkerPublishes result commitment
  8. →
  9. ClientVerifies the result
  10. →
  11. ChainPayment, or fraud proof and slash

The client checks the result with a Freivalds test. If one value is wrong, anyone can prove it on chain in linear time, the worker's bond is split between the challenger and the burn, and the client is refunded.

2.3Classes of work

Every block declares the class of work it contains. The class decides which matrices the validator recomputes against.

ClassMatrices come fromWho may mine itReward
NativeThe protocol: derived from the epoch number (100 blocks per epoch)Any minerBase reward. No usefulness is claimed.
JobA client's committed matrices A and BOnly the worker of that jobBase reward plus a capped rebate
ProductOne n × n block of a large tiled productOnly the worker of that jobSame as Job
OpOne GEMM inside a verified tensor programOnly the worker of that programSame as Job
LayerLegacy multi-layer inference jobs, superseded by programsOnly the workerSame as Job

Because the base reward is identical for every class, the market can be empty and the chain is exactly as secure as it would be without the market. The useful classes add a rebate that can never exceed what the client burned at creation.

03 · The proof of workA noised matrix product, sliced into tickets.

3.1Inputs and commitments

Work is defined over two n × n matrices A and B of signed 8-bit integers with every entry in [−64, 64]. A is committed by a Merkle tree over its rows and B by a Merkle tree over its columns. The roots are a_root and b_root; in a job they are fixed by the client's transaction. Each leaf is hashed with its index under a dedicated domain label, so a row can never be confused with a column or with a node.

3.2Low-rank noise

A per-attempt seed σ is derived from the parent block, the miner, the work class, an extra nonce, the transaction root and the timestamp. From σ, an extendable-output function produces four small matrices with entries in {−1, 0, 1}: EL and FL of size n × r, and ER and FR of size r × n. The miner multiplies the noised matrices:

// r ≤ 63, so every noised entry still fits in a signed 8-bit integer A′ = A + EL·ER B′ = B + FL·FR C′ = A′ · B′ // exact int32 arithmetic

The noise makes every attempt depend on a fresh seed, so no work can be reused from one block, miner or job to another. Because σ commits to the miner and to the transaction root, changing the beneficiary or the transactions invalidates the work Implemented.

3.3Attempts and the jackpot

Each 16 × 16 tile of C′ is one attempt. For tile (ti, tj) the accumulator is built by summing the inner dimension in chunks of fold_k steps. After each chunk, the 256 accumulator values are folded into a 64-word state by XOR with a step-dependent rotation. The final state is hashed once.

acc ← 0 ; fold[0..63] ← 0 for s in 0 … n/fold_k − 1: acc += Σk in chunk s A′[tile rows][k] · B′[k][tile cols] // mod 2³² rot = (s mod 31) + 1 for each of the 256 values i: fold[i mod 64] ⊕= rotl32(acc[i], rot) jackpot = BLAKE3( "vearl/jackpot/v2" ‖ σ ‖ ti ‖ tj ‖ fold ) wins if first 8 bytes of jackpot, read big-endian, < target

Folding is what lets a GPU mine at tensor-core speed. Hashing every partial sum would cost far more than the multiplication itself. Folding makes the hash cost negligible while still forcing the miner to compute every partial sum. Hypothesis The security of this folding step has no formal analysis. It is the most original and least tested component of the design, and it is the first place an auditor should look.

One mining attempt
  1. InputMatrices A, B and the seed σ
  2. →
  3. NoiseA′ = A + E·E, B′ = B + F·F
  4. →
  5. GPUint8 product, tile by tile, folded
  6. →
  7. HashBLAKE3 of the folded state
  8. →
  9. TestBelow target? The tile is a block

The noise is removed afterwards in O(n2r) operations, which recovers the exact product A·B that the client asked for.

3.4Verifying a block

A block carries a proof: the 16 rows of A and 16 columns of B that the winning tile needs, each with a Merkle path to the committed roots. A validator checks the paths, re-derives the noise from σ, recomputes that single tile, and compares the hash. The cost is O(162·n) instead of O(n3). Proofs weigh about 6 KiB at n = 128 and about 43 KiB at n = 1024 per header. That is the price of a useful proof of work that can be verified without the job data Measured.

3.5Decoding the useful result

The worker needs the exact product, not the noised one. The noise terms are low rank, so they are removed cheaply:

C = C′ − A·(FL·FR) − (EL·ER)·B − (EL·ER)·(FL·FR) = A·B

3.6A flaw we found, and the rule that closes it

A client that also works as the worker can commit matrices of very low rank, in the extreme case all zeros, which are valid for the protocol. Partial sums of such a product are much cheaper to compute, so each attempt costs less than honest work. The attacker gets more tickets per unit of energy and divides the cost of a majority attack. This applies to any useful proof of work where the client picks the inputs.

The attack is real: we implemented it and its jackpots are bit-identical to the honest computation. With the original parameters (rank 16, fold_k 256) an attempt cost 0.06 times the honest computation. The noise guarantees the effective rank of A′ and B′ is at least r, so the gain is about fold_k / r and vanishes when the folding interval does not exceed the noise rank Measured.

Noise rankfold_k 163264128256
161.000.500.250.130.06
322.021.010.500.250.13
634.002.001.000.500.25

Attacker cost divided by honest cost per attempt, all-zero inputs, n = 4096. Values of 1 or more mean the shortcut gives no advantage.

The genesis block of any chain is rejected unless its parameters pass a security check: tile 16, n a multiple of 128, 1 ≤ rank ≤ 63, fold_k a multiple of 32 that divides n, and fold_k ≤ 1.05 × rank. The testnet uses n = 128 (larger for model workloads), rank 63 and fold_k 64 Implemented.

Still open

Low rank is one structured family. Toeplitz, circulant, low-displacement-rank or block-sparse matrices admit fast products, and their effect on the cost of partial sums has not been analysed. The noise distribution {−1, 0, 1} was chosen to fit in 8 bits; its effect on the indistinguishability of tiles is not analysed either Open.

3.7GPU implementation and its real overhead

The CUDA kernel uses int8 tensor cores (mma.m16n8k32), asynchronous copies, folding in registers and an on-GPU BLAKE3. Its output is tested bit for bit against the CPU reference for jackpots, noise generation, exact products and mined blocks, so a node without a GPU verifies the same blocks. On an NVIDIA A10 with rank 63 and fold_k 64 Measured:

68–79 M

attempts per second, n = 128 to 512 (71 M at n = 1024)

≈ 9 %

consensus overhead versus the plain GEMM of the same kernel (n = 8192)

≈ 87

TOPS int8 reached by our kernel, where cuBLAS reaches about 199

The 9 % figure is relative to our own kernel. Our kernel is itself about 2.3 times slower than cuBLAS on the same card, so a miner who would otherwise use cuBLAS pays roughly 2.5 times in time today. Closing that gap is engineering work that has not been done. The overhead figures published by other projects are measured on other hardware and models and are not comparable. Only one GPU model (A10) has been measured; the kernel compiles for compute capability 8.0, 8.6, 8.9 and 9.0 and refuses older cards.

04 · Chain consensusBlocks, fork choice and difficulty.

4.1Accounts and transactions

The ledger is account based. An account is a 32-byte public key with a balance and a nonce. A transaction carries the sender, a sequential nonce, a fee and a typed body, and is signed over its canonical encoding. Fees are credited to the miner of the block that includes the transaction. A block containing a single invalid transaction is invalid as a whole.

TransactionPurpose
TransferMove funds between accounts
SubmitJob, SubmitJobAuctionCreate a job with an escrow, optionally with a reverse auction
Bid, ClaimJobOffer a lower price, then claim the job and post the bond
PostResultPublish the commitment to the result
Accept, SettleClient approval, or permissionless settlement once the challenge window closes
FraudProve that one value of a result is wrong
Fetch, Serve, SlashUnservedOn-chain data-availability challenge, response and penalty
Tiled and program jobsLarge products and verified tensor programs, with their own posting and fraud transactions

4.2Block header

FieldContent
version, height, prevHeader version (2), block number, hash of the parent header
timestampSeconds; strictly greater than the parent's
minerPublic key that receives the reward; bound into the work seed
class, extra_nonceClass of work and the free nonce a miner varies
ti, tj, jackpotThe winning tile and its hash
tx_rootMerkle root of the transactions (all zeros if none)
targetDifficulty target, 64 bits
a_root, b_rootRoots of the matrices the work was done on, so a tile proof can be checked without job data
accounts_rootMerkle root of all live accounts after the block, so a balance can be proven to a light client

4.3Validating a block

A validator applies these checks in order and rejects the block at the first failure.

  1. The version, height and parent hash are correct.
  2. The timestamp is strictly greater than the parent's.
  3. The target equals the target the difficulty rule prescribes.
  4. The announced jackpot is below the target. This cheap filter runs before any expensive check, which limits denial-of-service.
  5. The transaction root matches, the block has at most 256 transactions and at most 4 MiB of serialised transactions.
  6. The work is valid: the matrices follow from the class, the seed is recomputed, the tile is recomputed from the proof, and the jackpot matches.
  7. Transactions are applied in order against the account state.
  8. The block reward and any capped rebate are credited to the miner.
  9. The resulting account root matches the header.

A separate node-local rule refuses blocks whose timestamp is more than 60 seconds ahead of the local clock at the moment of reception. It is deliberately not part of consensus, so a chain that was valid when mined is still accepted when replayed later.

4.4Fork choice

Each block has a work value equal to 264 divided by (target + 1). The canonical chain is the valid chain with the greatest cumulative work, not the longest. A candidate replaces the current chain only if its cumulative work is strictly greater, so on a tie the node keeps the chain it saw first.

4.5Difficulty: ASERT

The target adjusts at every block with an exponential rule anchored to an ideal schedule Implemented Literature:

target = anchor_target · 2^( (Δt − T·(Δh + 1)) / halflife ) T = 30 s, halflife = 900 s (30 blocks), anchor = first block, integer arithmetic only

Because the target depends only on how far real time has drifted from the ideal schedule, there is no sliding window for an intermittent miner to game. Our first rule, LWMA-1 [12], could be slowed by a miner that switches on and off. We measured this by simulation Simulated:

RuleSteadyOn/off ×20After ×10 jumpAfter ÷10 drop
LWMA-1 (30 blocks)30.3 s38.3 s (+27.8 %)20.7 s58.6 s
ASERT, halflife 900 s30.0 s30.0 s (−0.1 %)7.8 s62.1 s

Average block time in seconds, steady hash rate and over the 100 blocks that follow a sudden change. The cost of ASERT is a temporarily faster emission after a sustained jump in hash rate, which the absolute schedule later claws back.

The implementation was checked against the 29 official test vectors of the Bitcoin Cash Node reference and against an independent Python implementation on 50,000 random inputs, with no difference Measured. A timestamp pushed 60 s into the future makes the next block at most about 4.7 % easier, and the next block corrects it. The simulation uses exponential solve times and no network latency, so it gives orders of magnitude, not guarantees.

4.6Reorganisations and synchronisation

  • A node keeps full-state snapshots every 16 blocks (the last 64 plus genesis) and can rewind to any height by replaying at most 15 blocks.
  • On a fork, it locates the common ancestor by stepping back 8, 16, 32 … blocks, downloads only the diverging branch, and performs an atomic reorganisation: if any block is invalid, the work is insufficient or the brake refuses, the original chain is restored exactly.
  • Transactions from the abandoned branch that the winning branch did not include return to the mempool.
  • Sync replies are bounded to 8 MiB, and the requester continues until it receives an empty batch.

A deep-fork test merged two chains of about 155 blocks with jobs and transfers on each side in a few seconds, with the supply invariant intact Measured.

4.7Optional deep-reorganisation brake

Vearl includes an opt-in brake modelled on ECIP-1100 [9]. A competing chain is adopted only if its work since the common ancestor exceeds f(t) times ours, where t is the time our chain has been in place and f rises smoothly from 1 to 31 over about 7 hours. It is off by default. A more aggressive first setting made two equal halves separated for only 90 seconds unable to rejoin, because each demanded about 5 % extra work from the other. Liveness matters more on a young network, so operators enable the brake once partitions longer than a few minutes are unlikely. It is a subjective local rule, and a long partition then needs manual recovery.

In simulation, with the brake on, an attacker holding 45 % of the hash power has a 0.4 % chance of replacing 30 or more blocks within 4 hours, against 40.5 % without it Simulated.

4.8What consensus does not provide

  • No finality. Reorganisations are braked, not impossible.
  • No protection against selfish mining. In simulation, a selfish miner's revenue share is 0.274 with 30 % of the hash power (no gain), 0.327 with 33 % and 0.387 with 36 %, so the profitable threshold is about one third. As in any Nakamoto chain, this is inherent [11] Open.
  • No protection against a majority of the work, and a very small security budget at launch (section 10).

4.9Light clients

A light client downloads headers, each with its tile proof, and verifies the chain, the timestamps, the difficulty, the native work roots and the recomputed tile without any state. It then checks an account balance with a Merkle path against accounts_root of the tip. Tested on a live node, 85 headers were verified in 0.2 s and a balance was proven Measured. As with any SPV client, transaction validity and the honesty of accounts_root rest on the honest majority of the work, and proofs of account absence do not exist yet.

4.10Monetary invariant

After every block the node checks that the sum of balances, open escrows, bonds, held data-request fees and everything ever burned equals the genesis allocation plus everything ever minted. A violation would indicate a consensus bug. This invariant is verified in tests, in fuzzing and in the independent implementations Implemented.

05 · The compute marketEscrow, bonds and fraud proofs.

The market is part of the protocol, not an application on top of it. Every rule below is enforced by validators.

5.1Job lifecycle

State machine of a job
  1. StateOpen escrow locked, 10 % burned
  2. →
  3. StateClaimed worker bond posted
  4. →
  5. StatePosted result committed, challenge window open
  6. →
  7. EndSettled worker paid
  1. Posted →Slashed fraud or unserved data
  2. Open or Claimed →Expired client refunded

Every transition is evaluated with the height and timestamp of the block that contains the transaction.

StepRule
SubmitEscrow of at least 1 unit. 10 % is burned immediately. The remaining 90 % is the net escrow. The burned amount also becomes the job's rebate budget. Commitments a_root and b_root are fixed here.
ClaimThe worker posts a bond of 10 % of the agreed price. With an auction, the winner claims at its own bid and the client immediately recovers the difference.
PostThe worker publishes c_root, the Merkle commitment to the result rows. The challenge window starts.
WindowStays open until both 10 blocks and 300 seconds of timestamp have passed.
Accept / SettleThe client can accept early. Otherwise anyone can settle after the window. The worker receives the net escrow plus its bond back.
FraudA valid proof slashes the bond: half to the challenger, half burned. The client receives the net escrow back.
ExpireAn unclaimed job refunds the client after the deadline. A claimed job with no result burns the bond and refunds the client.

5.2Why the window counts both blocks and seconds

A design flaw found during testing: the window was first measured in blocks only. On a fast network, or with a miner producing blocks back to back, ten blocks pass in under a second, and a cheating worker was paid before the client could verify anything. The window now also requires 300 seconds of timestamp, which a burst of blocks cannot consume because timestamps are bounded by validators' clocks. A test confirms that 40 blocks one second apart do not close the window Implemented.

5.3Reverse auctions and reputation

A job may carry an auction. During it, workers submit bids that must be strictly lower than the best one and no higher than the net escrow. Nobody can claim before it ends. The winner has 60 seconds of exclusivity, after which any worker can claim at the full price, so a winner that disappears loses its place without blocking the job. Bids are public and sequential, so last-second undercutting is possible Open. Workers accumulate counters of jobs done, slashed and expired. They are informational and not enforced by the protocol.

5.4Data availability

A worker could publish a commitment without holding the data, or hold it back after being paid. During the window, the client can demand any single entry of the result on chain by paying a small fee (Fetch, up to 8 pending at once). The worker must answer with the value and a Merkle path within 120 seconds (Serve); answering earns the fee and reopens the window so the client can check the value. If it does not answer, anyone can submit SlashUnserved: the reporter receives half of the bond, half is burned, and the client is refunded. Automatic settlement is blocked while a request is pending.

This guarantees that nobody is paid for a commitment with no data behind it, and that a fabricated result is caught by a single sampled entry. It does not deliver the full result on chain: delivery of a correct result that a worker refuses to hand over remains off chain, and fair exchange without a trusted third party is a known hard problem Open.

5.5Usefulness and the rebate

A block of a useful class earns the base reward plus a rebate of min(base / 2, rebate left), where the rebate budget starts at the amount burned when the job was created. Usefulness is therefore subsidised by at most what clients burned. Consider a wash trader who is both client and worker of its own job. It pays the 10 % burn and can recover at most that amount as rebate, so it never earns more than mining native work, and a test asserts this Implemented. The 10 % burn is therefore not a revenue mechanism or a deflationary policy. Its role is to bound the rebate.

06 · Verifying computationExact where it can be, honest where it cannot.

6.1Three levels of guarantee

LevelVerificationMines at the same time?Guarantee
1 · Verifiable GEMMFreivalds test plus single-entry fraud proofsYes: the client's noised product is the consensus workStrong and cheap. One wrong entry is always provable.
2 · SamplingRecompute random chunks, fraud proof on one chunkNo. Paid at full cost, without rebateProbabilistic. An isolated falsified chunk is missed with probability 1 − m/N for m checked chunks of N.
3 · Not cheaply verifiableNoneNoNone without trust. Not supported.

Most of this paper concerns level 1, which is where the protocol is strongest. Level 2 is implemented for two fixed kernels and is described in section 6.7.

6.2Freivalds check

To check C = A·B, pick a random vector x and compare A·(B·x) with C·x. Each round costs three matrix-vector products, O(n2), and a wrong result survives a round with probability at most 1/2. The client runs 24 rounds, for an error of at most 2−24 Literature [4].

6.3Single-entry fraud proofs

Freivalds tells a client that something is wrong. To convince the chain, the client locates one wrong entry (i, j) and submits:

  • row i of A with its Merkle path to a_root,
  • column j of B with its path to b_root,
  • row i of the published result with its path to c_root.

Validators verify three paths and compute one dot product. The proof is valid, and the worker is slashed, if and only if the dot product differs from the published entry. It costs O(n) and is about 1.2 KB at n = 128. An honest worker cannot be slashed, because the proof of a correct entry fails.

Integer matrices are committed by rows: one leaf per row, written in little-endian int32. An earlier version used one leaf per entry, which required n2 hashes instead of n. At n = 1024, that was about 100 times slower (396 ms against 4 ms on one core). Switching to rows cut client verification of a 76-operation network from 130 s to 18 s for 32 images Measured.

6.4Verified tensor programs

A single product is not a neural network. A program is a sequence of at most 240 operations on n × n int8 tensors with at most 64 weight blocks. Tensor T0 is the client's input and operation k produces Tk+1. Every tensor is committed by columns, weights by rows, and the raw int32 result of each GEMM by rows. The worker publishes the roots as it goes, and each operation has its own fraud proof, so an error anywhere is provable on its own. Later operations computed from a wrong value are correct relative to it.

OperationDefinition (all integer, all fixed)Proof shape
Gemm, GemmATT = clamp((W·Tx) ≫ shift), or the product of two computed tensors (attention scores). The mined work.Product and activation
AddSaturating sum; residual connectionsElement
MapTranspose, im2col and general strided convolution, flatten, shift, multi-head spread and gatherElement
PoolAverage (rounded) or max over k × k windowsElement
Softmax, SoftmaxLoInteger softmax per column from a fixed exp table, plus a fine part giving probabilities to 1/4096Column
CausalSoftmaxThe same, where query q only sees keys 0 … q; used by language modelsColumn
LayerNormRounded integer mean and standard deviation, exact integer square rootColumn
Bias, Mul, LutPer-row bias, per-row multiply, and a 128-entry lookup table (GELU, tanh, SiLU and others)Column, row, table

Activations saturate to [−64, 64], shifts are arithmetic, and no operation uses floating point. The exponential table is computed offline in exact decimal arithmetic. Two honest workers therefore always produce identical bits. The second stage of a large softmax (SoftmaxLo) was needed because 26 to 37 % of the attention mass of a 197-token image sits in probabilities below 1/128, which a one-byte softmax rounds to zero.

6.5Models, quantisation and chaining

A floating-point model is not the same model once quantised to 8 bits, and the gap must be measured rather than assumed. Each model is converted to a program, then fine-tuned with quantisation-aware training in a simulator that uses exactly the integer arithmetic of the chain. The simulator is checked bit for bit against the independent reference engine before and after training. A job holds at most 240 operations, so larger models are split into jobs that chain: the verified output of one is the input of the next, for example three jobs of four layers for a 12-layer encoder.

6.6Tiled products

A product larger than one block (up to 8 × 8 × 8 blocks) is split into block products, each a mined unit, plus verified sums. A block product or a sum can each be proven wrong with a proof of the same shape. No block product is required to advance the chain: any winning attempt from any block of the job can become a block. The client verifies with Freivalds on each block and exact sums. An 8192 × 8192 × 8192 product (about 1.1 trillion integer operations) was verified end to end in about 11 s Measured.

6.7Level 2: sampled audits

Monte Carlo integration and tiled 3D rendering run as protocol-fixed integer kernels, with no client code executed and no sandbox. Each chunk is a deterministic function of public parameters and its index, and the worker commits to the Merkle root of the chunks. The client recomputes a random sample; one fraud proof names a chunk, which the chain recomputes at bounded cost. As stated in the table above, the guarantee is probabilistic and these jobs carry no rebate. A Monte Carlo estimate of π with 262,144 samples gave 3.14497 Measured.

07 · Verified AI workloadsWhat runs today, with the numbers.

Every figure below was measured on the testnet on a single NVIDIA A10 GPU, on chains with n = 1024, with every operation of every job verified before payment. Outputs were compared with an independent reference engine that shares no code with the node.

WorkloadModelSizeMeasured
Text generationTinyStories-33M, 4 layers, 50,257-word vocabulary [15]2 jobs per word (154 + 153 ops)Eight stories in parallel, 320 words in 730 s (18 s per step). The next word equals the original model's in 92.9 % of positions; perplexity 2.42 against 2.37. Logits bit-identical to the reference engine, 8 of 8.
Photo recognitionDeiT-Tiny vision transformer, 224 × 224, 1,000 classes [13]386 ops, 2 chained jobsAbout 1.7 s per photo. Top-1 over 1,000 classes: 70.6 % on Imagenette photos not seen in training (original: 78.2 %). 8 of 8 bit-identical.
Text embeddingsBGE-base, 12 layers, 110 million parameters [14]492 ops, 3 chained jobsAbout 3.3 s per text. STS-B Spearman 0.836 against 0.864 for the original (96.7 %). 8 of 8 bit-identical.
Fast embeddingsall-MiniLM-L6-v2, 6 layers [16]195 ops, 61 GEMMsAbout 1.0 s per text; cosine 0.92 with the original.
Image classificationResNet-20, CIFAR-10 [17]76 ops, 22 GEMMs88.6 % top-1; 32 images verified in 41 s.
Linear algebraTiled int8 product8192 × 8192 × 8192About 1.1 trillion operations verified in about 11 s.
Digits, bulk embeddingsMLP and CNN on MNIST; distilled text encoder10,000 images; 20,480 textsAbout 98 % top-1; outputs bit-identical to the reference engine.

7.1How to read these results

  • Integer models lose accuracy. DeiT-Tiny is about 8 points under its floating-point original. The language model is slightly worse than the original but writes coherent stories. Models whose activations have very large outliers fail outright: a simulated 8-bit trial on GPT-2 (124 M parameters), without training and after our best exact transformations, gave a perplexity of 1,345 against 99, which is why the language model on the testnet is small and trained on simple text.
  • Language-model exactness. The chain's words follow the original model word for word over 68 of 320 generated words, then diverge. That is expected with greedy decoding when two models differ slightly, and the stories remain coherent. What is exact is the computation of each step, not agreement with the floating-point original.
  • Limits. Text is capped at 64 tokens. There is no key-value cache, so each step recomputes the whole text. A job has 240 operations, which reaches models of roughly a hundred million parameters, not billions.
  • Verification is slower than running locally. These times include transfer and cryptographic commitment. When client and worker share a machine, local computation is faster. The value of the chain is verified work done by others, not speed.

08 · CryptographyStandard pieces, nothing homemade.

PurposePrimitive
Hashing, Merkle trees, key derivationBLAKE3 [7] in key-derivation mode with a distinct context string for every use (vearl/merkle/leaf/v1, vearl/jackpot/v2, and so on)
Classic account signaturesEd25519 [10], strict verification
Post-quantum account signaturesSLH-DSA-SHAKE-128s, FIPS 205 [5]
Peer channelNoise XX with X25519, ChaCha20-Poly1305 and BLAKE2s [6]
Wallet fileArgon2id and ChaCha20-Poly1305 (RFC 9106, RFC 8439), 24-word recovery phrase
AddressesBech32m (BIP-350): vrl1… classic, vrlq1… post-quantum

8.1Post-quantum accounts

An account can be classic or post-quantum, and the protocol accepts both. The signature length selects the scheme: 64 bytes is Ed25519, 7,856 bytes is SLH-DSA, and any other length is refused. The 32-byte public key stays the address in both cases, so no transaction format changed. The signed message uses a dedicated FIPS 205 context string, so a signature made in another context is refused. A post-quantum key is derived deterministically from the wallet phrase.

Ed25519SLH-DSA-SHAKE-128s
Signature64 bytes7,856 bytes
Signing timemicroseconds≈ 1.6 s
Verificationmicroseconds≈ 1.6 ms
Security assumptionDiscrete logarithm (broken by Shor's algorithm)Hash functions only

An independent verifier written from the FIPS 205 text agreed with the Rust crate on 16 of 16 verdicts, including 14 tampered signatures Measured. Limits: the Rust crate is pure Rust and uses no unsafe code but has not been audited; there is no fee schedule per byte yet, with the cap of 256 transactions per block bounding the worst case at about 2 MiB; and the peer channel is not post-quantum. A classic public key that has already been published stays exposed to a future quantum computer until funds move to a post-quantum account.

09 · NetworkEncrypted peers, bounded everything.

  • Transport. TCP carrying a Noise XX channel. A peer's identity is its static key, and operators can pin it (host:port@key). Plaintext peers of older versions are refused. An application message is at most 16 MiB.
  • Handshake. Each side sends Hello first with its genesis hash, height and cumulative work. A different genesis closes the connection, and so does an identity that is already connected.
  • Admission. At most 24 inbound connections, 2 per IP and 4 per /16 subnet; at most 6 outbound connections, never two in the same /16; 64 concurrent handshakes. This raises the cost of an eclipse attack. It is not a Sybil defence Open.
  • Rate limiting. A token bucket of 100 messages per second per peer, with a burst of 200. Invalid messages accumulate a score, and an invalid block or 100 points disconnects and bans the address.
  • Mempool (a local rule, not consensus). Signature and nonce checks, a non-empty balance, 16 transactions per sender and 5,000 in total. A sender whose first pending transaction is stuck, for 10 s on a nonce gap or 30 s on failure, has all of its pending transactions evicted. Without this rule a single stuck transaction froze a worker completely; the endurance test found it.
  • Data. Clients upload job data to a worker off chain. The worker persists it (up to 512 MiB) and reloads it on restart. Intermediate results are fetched in bounding boxes so that sparse tensors transfer compactly.

The P2P channel is not post-quantum, and IP addresses of peers are visible: there is no network-level anonymity.

10 · EconomicsProvisional numbers, clearly labelled.

Read first

Nothing in this section predicts a price, a demand or a profit. The parameters are testnet values chosen to exercise the mechanisms. They have not been calibrated with market data. Testnet tokens have no value.

10.1Block reward and emission

The reward of a block of height h is max(R0 ≫ ⌊h / 1,051,200⌋, tail) with R0 = 20 units and a tail of 0.5 units per block. At the target block time of 30 seconds, 1,051,200 blocks correspond to one year. Halving depends on block height, not on time. The schedule below is derived directly from these constants.

YearReward per blockEmitted in the yearCumulative
12021,024,00021,024,000
21010,512,00031,536,000
355,256,00036,792,000
42.52,628,00039,420,000
51.251,314,00040,734,000
60.625657,00041,391,000
7 onwards0.5 (tail)525,600 each yearno cap

Units, before any genesis allocation. The supply is not capped: the tail keeps a small permanent issuance to pay for security.

10.2Who pays whom

  • Security is paid by the block reward, identical for every class of work.
  • Usefulness is paid by clients. A rebate on useful blocks is bounded by what the client burned.
  • Workers earn block rewards for the tickets their computation produces, and the escrow for results verified. In a competitive equilibrium a client's job costs little more than the duplex overhead when demand is small compared with issuance. The issuance then subsidises the computation Simulated.
  • That subsidy means the token needs a value independent of usage. A fixed minimum escrow in the token is arbitrary next to a product that costs a tiny fraction of a cent. An adaptive floor, similar to a fee market, is planned.

10.3The security budget is small at launch

The cost of an attack is tied to issuance times price. In a model where a rented GPU costs 0.50 dollars per hour, a hypothetical token price of 0.10 dollars attracts about 1,455 GPUs to compete for the issuance, so matching the honest network costs an attacker about 730 dollars per hour. A value of 10,000 dollars would then only be safe after about 14 hours of confirmations Simulated. This is the standard young-chain problem, and no parameter creates security without economic value behind the token. A network that carries real value should only accept amounts below the cost of attack over the confirmation time.

10.4What is not modelled

Strategic behaviour (selfish mining, withholding, eclipse), endogenous price volatility, bandwidth and storage costs, real demand, competition from other chains for the same GPUs, and regulation. A token with monetary value would need a legal review before anything is offered to anyone. None of that has been done, and no sale of any token is planned or implied by this paper.

11 · Security analysisWhat is covered, what is not.

The threat model considers a hostile peer, a miner trying to get blocks with less work than honest, a worker that cheats or withholds data, a client that accuses falsely, a Sybil or eclipse attacker, and an attacker with large computing power. Protection against a majority of the work is not part of the model. Client-side compromise (malware, phishing) is out of scope.

ThreatDefenceStatus
Cheap attempts through structured, low-rank inputsGenesis parameter rule, fold_k ≤ 1.05 · rankCovered for the parameters accepted; other structures are a hypothesis
Work reuse across blocks, miners, jobsThe seed binds parent, miner, class, transactions and timeCovered
Forged tile proofMerkle openings plus recomputation of the tileCovered
Wrong job result24-round Freivalds check and a single-entry fraud proof with a bondCovered
Being paid before the challenge windowWindow in blocks and in secondsCovered
Withholding result dataFetch / Serve / SlashUnservedPartial: full delivery stays off chain
Wash trading to inflate rewardsRebate bounded by burnCovered by design
Difficulty manipulationASERT, timestamp rulesCovered by simulation, not proof
Malformed input crashing a nodeBounded decoding, errors instead of panics, fuzzingCovered empirically
Supply inflation bugsInvariant checked after every blockCovered empirically
Node diverging from the specificationByte-level spec, independent Python implementationsCovered for the stated scope; same author
Eclipse and SybilPer-IP and per-/16 limits, pinned keysPartial, no Sybil defence
Deep reorganisation, majority attackOptional brakePartial; short attacks and selfish mining untreated
Level 2 inexact resultSampling auditsPartial by design: probabilistic
Confidentiality of data and transactionsNone. Everything is public or visible to the workerOpen
Errors shared by code, text and second implementationNoneOpen: a third-party implementer is needed
Soundness of the folding stepEmpirical study onlyOpen
Reminder

Vearl makes no claim to be resistant to ASICs. An int8 matrix-multiplication kernel sits exactly where accelerators already exist. The main defence is real demand, since the same GPUs also serve AI workloads, and not an algorithm. The design does not claim that no attack exists, total anonymity, or that it is better than other proofs of useful work in practice. Its differences are design differences, and superiority must be shown by measurements.

12 · Implementation and evidenceChecked against a second opinion.

The reference implementation is written in Rust, with a CUDA kernel for the GPU. It contains the node, the wallet, a command-line client, an HTTP API with a block explorer, a faucet, and deployment units. The code has not been published yet. This paper therefore states which claims are backed by which checks, so that reviewers can re-run them once they have access.

171

automated tests, all passing

17 / 17

independent checks in the internal audit suite

483

conformance blocks shared by Rust and Python

0

external audits

12.1Independent implementations

The specification describes every byte of the work layer, the chain, the market, the tensor programs and the network. Separate Python implementations were written from the text, without reading the Rust code, for the work layer, the chain, the market, tensor programs, tiled products, the light client and the post-quantum verifier. Each one revalidates the frozen test vectors and refuses the tampered ones for the right reason. They also mine and author blocks that the Rust validator must accept, in both directions. A conformance run of 2,994 random tensor programs compared every tensor with the node's engine and found no divergence, and a differential fuzzer showed 1,500 of 1,500 identical accept or reject decisions between Rust and Python.

12.2Robustness

  • About 1.1 million mutated transactions and about 460,000 mutated blocks and messages produced no panic and no accepted altered block.
  • A network fuzzer sent about 18,000 malformed messages to a live worker node: no panic, the node stayed reachable and kept mining.
  • A chaos test with four real nodes, abrupt kills, jobs of every type and a cheating worker ended with all honest jobs paid and all cheaters slashed.
  • The Python and Rust difficulty implementations agreed on 50,000 random cases and on the official ASERT vectors.

12.3A limit of this evidence

The code, the specification and the second implementations were all written by the same party. A truly independent third-party implementer and an external audit are still needed. The checks above show that the pieces agree with each other, not that they are free of a mistake that they would all make together Open.

13 · LimitationsWhat this paper does not claim.

  1. Useful means paid and burned, not proven useful. A client can pay for useless computation at the price of the burn. This trade-off exists in every useful proof of work where the client chooses the input.
  2. The folding step has no formal security proof, and structured inputs other than low rank have not been analysed.
  3. Freivalds proves a product is correct, not that a particular GPU computed it. That suffices to pay for a result. Proof of hardware work comes from the lottery.
  4. Data availability is partial. Complete delivery of a correct result stays off chain.
  5. One worker per job, first come, first served, unless an auction is used, and auction bids are public.
  6. No confidentiality. The chain is public and workers see their clients' matrices and weights.
  7. No finality, no defence against selfish mining, and a very small security budget at launch.
  8. Models are small and lossy. A job has 240 operations on 1024 × 1024 tensors. Quantisation to 8 bits costs accuracy. Language-model text is capped at 64 tokens and computed without a cache. These are first-generation models, and they are planned to be considerably improved (see the milestones).
  9. Only one GPU model has been measured, and the kernel is well under the speed of cuBLAS.
  10. Economic parameters are provisional. Nothing has been calibrated on real demand, and the minimum escrow is badly specified.
  11. No external audit, and no public testnet yet (planned for October to November 2026).

14 · Status and roadmapWhere the project stands.

Milestones
Done
Closed testnet. Chain, compute market, verified inference, embeddings and text generation on a GPU server, with the light client, post-quantum accounts, wallet, API and explorer.
Oct – Nov 2026
Public testnet. Several machines, external miners, and open documentation. Planned for October to November 2026. Miners who help secure the public testnet are planned to receive Vearl tokens on mainnet.
Next
External audit. Third-party review of the cryptography, the folding step, the consensus rules and the fraud proofs, and an independent second implementation.
Research
Considerably better models, and better economics. The current models are a first generation; larger and more accurate ones are planned through per-channel quantisation, longer text and a key-value cache, several sequences per job, an on-chain model registry, an adaptive escrow floor, a fee schedule per byte, confidentiality, and a faster kernel.

A mainnet should exist only if the security properties above have been demonstrated and reviewed. There is no date for it, and the testnet chain can be reset at any time. Miners who help secure the public testnet are planned to receive Vearl tokens on mainnet as a thank-you; this is an intention, not a guarantee or an offer, and the rules and amounts will be announced.

Feedback. This is a living document (draft 0.1). Where it differs from the behaviour of the software, the software on the testnet is the reference, and we want to hear about it. Corrections to numbers and claims are welcome, especially from reviewers with a security or distributed-systems background.

Nothing in this paper is an offer to sell anything or investment advice. Testnet tokens have no value.

Appendix AParameters and constants.

Values of the current testnet. Those marked genesis are part of the genesis hash, and changing them creates a different chain.

ParameterValueNote
Work parameters (genesis)n 128, tile 16, rank 63, fold_k 64Model workloads use chains with n = 1024
Entry range[−64, 64]Signed 8-bit integers
Target block time30 s
Native work epoch100 blocksNative matrices change every epoch
Difficulty rule (genesis)ASERT, halflife 900 sLWMA-1, EMA and fixed also exist
Future timestamp tolerance60 sLocal rule
Block limits256 transactions, 4 MiBConsensus
Initial reward, halving, tail20 · 1,051,200 blocks · 0.5Units
Unit108 base units
Job burn, worker bond10 %, 10 %Provisional
Minimum escrow1 unitProvisional; to be made adaptive
Challenge window (genesis)10 blocks and 300 sBoth must elapse
Serve deadline (genesis)120 sWindow reopens for at least 60 s
Data request fee, pending requests0.05 unit, 8
Auctionup to 3,600 s; 60 s winner grace
Tensor program240 operations, 64 weight blocks
Snapshotsevery 16 blocks, last 64
Peer limits24 in, 6 out, 2 per IP, 4 per /16Local rule
Message size, rate16 MiB, 100 per s
Mempool5,000 total, 16 per senderLocal rule

Appendix BEncodings in brief.

The byte-level specification is long, and only its outline is reproduced here. All integers are little-endian, enumerations carry a 32-bit tag, and structures use the bincode 1.x layout. The Merkle padding repeats zero leaves up to the next power of two.

ItemDefinition
HashH(domain, p1, p2, …) = BLAKE3 derive-key with context domain of p1 ‖ p2 ‖ …
Merkle leaf, nodeH("vearl/merkle/leaf/v1", i, data); H("vearl/merkle/node/v1", left, right)
Header hashH("vearl/header/v1", bincode(header))
Work seed σH("vearl/sigma/v1", prev, miner, class, extra_nonce, tx_root, timestamp)
NoiseXOF("vearl/noise/v1", σ, 4·n·rank) bytes, each b ↦ (b mod 3) − 1
Native matricesXOF("vearl/native/v1", epoch, 2·n²) bytes, each mapped to (b mod 129) − 64
Job identifierH("vearl/job/v1", client, nonce, a_root, b_root)
Work class tagsNative 0, Job 1, Layer 2, Product 3, Op 4
Transaction tagsTransfer 0, SubmitJob 1, ClaimJob 2, PostResult 3, Accept 4, Fraud 5, Settle 6, SubmitJobAuction 7, Bid 8, Fetch 9, Serve 10, SlashUnserved 11, tiled and program jobs from 12
Operation tagsGemm 0, Add 1, Map 2, Pool 3, GemmAT 4, Softmax 5, LayerNorm 6, Bias 7, Lut 8, Mul 9, SoftmaxLo 10, CausalSoftmax 11
SignatureEd25519 (64 bytes) or SLH-DSA (7,856 bytes) over bincode(transaction)

ReferencesSources cited.

  1. I. Komargodski, I. Schen, O. Weinstein. Proofs of Useful Work from Arbitrary Matrix Multiplication. arXiv:2504.09971; IACR ePrint 2025/685.
  2. The Usefulness Gap in Proof-of-Useful-Work. arXiv:2606.04819, June 2026. A single third-party study, cited for its argument; its measurements were not reproduced here.
  3. R. Pass. The Economics of Proof-of-Useful-Work. arXiv:2606.06700.
  4. R. Freivalds. Probabilistic machines can use less running time. IFIP Congress, 1977.
  5. NIST. FIPS 205: Stateless Hash-Based Digital Signature Standard. 2024.
  6. T. Perrin. The Noise Protocol Framework. noiseprotocol.org.
  7. J. O'Connor, J.-P. Aumasson, S. Neves, Z. Wilcox-O'Hearn. BLAKE3: one function, fast everywhere.
  8. Bitcoin Cash Node. aserti3-2d difficulty adjustment algorithm (ASERT), specification and test vectors, 2020.
  9. Ethereum Classic. ECIP-1100: Modified Exponential Subjective Scoring (MESS).
  10. D. J. Bernstein et al. Ed25519; IETF RFC 8032. See also RFC 9106 (Argon2), RFC 8439 (ChaCha20-Poly1305) and BIP-350 (Bech32m).
  11. I. Eyal, E. G. Sirer. Majority is not enough: Bitcoin mining is vulnerable. 2014.
  12. Zawy12. LWMA difficulty algorithm. GitHub.
  13. H. Touvron et al. Training data-efficient image transformers and distillation through attention (DeiT). arXiv:2012.12877.
  14. S. Xiao et al. C-Pack: Packaged Resources to Advance General Chinese Embedding (BGE). arXiv:2309.07597.
  15. R. Eldan, Y. Li. TinyStories: How Small Can Language Models Be and Still Speak Coherent English? arXiv:2305.07759.
  16. W. Wang et al. MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers. arXiv:2002.10957.
  17. K. He et al. Deep Residual Learning for Image Recognition. arXiv:1512.03385.
↑