SHEET 04
Equipment schedule
21 marksOpen weights, run on our own accelerators. Every rate below is derived from the cost model on sheet 02 — none of it is a list price we chose.
| Mark | Model | Parameters | Licence | Context | Replica | Throughput | Cost | List | Per |
|---|---|---|---|---|---|---|---|---|---|
| H-01CodeSoftware engineering agents | |||||||||
| M-01 | Qwen3-Coder-480B-A35BDefault planner and editor for the code agent. | 480B / 35B act | Apache-2.0 | 262.144k | 4 GPU | 3.2k tok/s | $0.531 | $1.06 | 1M tok |
| M-02 | DeepSeek-V3.2Long-horizon reasoning; used for repository-wide refactors. | 685B / 37B act | MIT | 163.84k | 8 GPU | 5.4k tok/s | $0.629 | $1.26 | 1M tok |
| M-03 | Llama-4-MaverickWhole-codebase context without chunking. | 400B / 17B act | Llama 4 Community | 1,000k | 8 GPU | 6.1k tok/s | $0.557 | $1.11 | 1M tok |
| M-04 | Kimi-K2 (1T-A32B)Strongest open agentic tool-use; reserved for hard runs. | 1000B / 32B act | Modified MIT | 131.072k | 16 GPU | 5.2k tok/s | $1.31 | $2.61 | 1M tok |
| M-05 | GLM-4.6Balanced cost/quality workhorse. | 355B / 32B act | MIT | 200k | 8 GPU | 4.4k tok/s | $0.772 | $1.54 | 1M tok |
| M-06 | Qwen3-32BDense, latency-sensitive edits and completions. | 32B | Apache-2.0 | 131.072k | 2 GPU | 2.4k tok/s | $0.354 | $0.707 | 1M tok |
| M-07 | Devstral-SmallSingle-GPU tier for cheap loops and CI checks. | 24B | Apache-2.0 | 131.072k | 1 GPU | 1.9k tok/s | $0.223 | $0.447 | 1M tok |
| H-02Biological designProtein language models | |||||||||
| M-08 | ESM-C 6BRepresentation model — embeddings and variant effect scoring. | 6B | EvolutionaryScale Community | — | 1 GPU | 15.0k res/s | $0.028 | $0.057 | 1M res |
| M-09 | ESM3-open 1.4BGenerative PLM — sequence, structure and function conditioning. | 1.4B | EvolutionaryScale Community | — | 1 GPU | 4.0k res/s | $0.106 | $0.212 | 1M res |
| M-10 | ProGen2-xlargeFamily-conditioned de novo sequence generation. | 6.4B | Apache-2.0 | — | 1 GPU | 1.5k res/s | $0.283 | $0.566 | 1M res |
| M-11 | Boltz-2Co-folding and binding affinity in one pass. | 1.2B | MIT | — | 1 GPU | 0.05 struct/s | $8.49 | $16.98 | 1k struct |
| M-12 | RFdiffusionBackbone design for binders and scaffolds. | 60M | BSD-3-Clause | — | 1 GPU | 0.2 backbone/s | $2.12 | $4.24 | 1k backbone |
| H-03VisionVision-language models | |||||||||
| M-13 | Qwen3-VL-235B-A22BFrontier-class open VLM; documents, video, GUI grounding. | 235B / 22B act | Apache-2.0 | 262.144k | 8 GPU | 3.9k tok/s | $0.871 | $1.74 | 1M tok |
| M-14 | InternVL3-78BHigh-resolution tiling for dense diagrams and scans. | 78B | MIT | 32.768k | 4 GPU | 2.2k tok/s | $0.772 | $1.54 | 1M tok |
| M-15 | Molmo-72BPointing and spatial grounding with open training data. | 72B | Apache-2.0 | 32.768k | 4 GPU | 2.0k tok/s | $0.849 | $1.70 | 1M tok |
| M-16 | SAM 3Promptable segmentation and tracking. | 900M | SAM Licence | — | 1 GPU | 30 img/s | $0.014 | $0.028 | 1k img |
| M-17 | dots.ocrLayout-preserving document parsing to structured text. | 1.7B | MIT | — | 1 GPU | 2.5 page/s | $0.170 | $0.340 | 1k page |
| H-04KnowledgeRetrieval, embedding, synthesis | |||||||||
| M-18 | Qwen3-Embedding-8BMultilingual retrieval embeddings. | 8B | Apache-2.0 | 32.768k | 1 GPU | 96.0k tok/s | $0.004 | $0.009 | 1M tok |
| M-19 | BGE-M3Dense, sparse and multi-vector retrieval in one pass. | 570M | MIT | 8.192k | 1 GPU | 210.0k tok/s | $0.002 | $0.004 | 1M tok |
| M-20 | Qwen3-Reranker-4BCross-encoder rerank over retrieved candidates. | 4B | Apache-2.0 | 32.768k | 1 GPU | 64.0k tok/s | $0.007 | $0.013 | 1M tok |
| M-21 | Llama-4-ScoutTen-million-token context for whole-corpus synthesis. | 109B / 17B act | Llama 4 Community | 10,000k | 4 GPU | 3.4k tok/s | $0.499 | $0.999 | 1M tok |
21 models in service. Input billed at 25% of the listed output rate. Licences are as published by each model's authors — check them before commercial use. Prices are derived from the cost model in note 02.1 and change when it changes.