Prime Intellect is an open-source decentralised AI research organisation and GPU compute marketplace founded by Vincent Weisser (CEO, formerly co-founder of VitaDAO and AI lead at Molecule biopharma) and Johannes Hagemann headquartered in San Francisco, operating on 4M total funding across three …
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:hasPart ai:PRIMEFramework))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:hasPart ai:DiLoCo))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:hasPart ai:ElasticDeviceMesh))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:hasPart ai:PRIMERRL))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:hasPart ai:TOPLOC))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:hasPart ai:SHARDCAST))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:hasPart ai:GPUMarketplace))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:hasPart ai:ComputeExchange))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:hasPart ai:INTELLECT1Model))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:hasPart ai:INTELLECT2Model))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:hasPart ai:INTELLECT3Model))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:hasPart ai:METAGENE1Model))Dependency Relationships
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:requires ai:GPUCluster))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:requires ai:DistributedTraining))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:requires ai:FaultTolerance))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:requires ai:LowBandwidthNetworking))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:requires ai:GradientSynchronisation))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:requires ai:CheckpointRecovery))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:dependsOn ai:PyTorch))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:dependsOn ai:FSDP2))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:dependsOn ai:HivemindLibrary))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:dependsOn ai:vLLM))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:dependsOn ai:libp2p))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:dependsOn ai:InfiniBandNetworking))Capability Relationships
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:enables ai:DecentralisedFoundationModelTraining))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:enables ai:PermissionlessComputeContribution))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:enables ai:GloballyDistributedReinforcementLearning))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:enables ai:OpenScienceAI))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:enables ai:ComputeDemocratisation))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:supports ai:LargeLanguageModels))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:supports ai:ReasoningModels))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:supports ai:MixtureOfExperts))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:supports ai:MetagenomicAI))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:supports ai:AgenticReinforcementLearning))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:supports ai:PandemicMonitoring))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:enables ai:CollaborativeModelOwnership))Implementation Relationships
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:implements ai:DiLoCo))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:implements ai:GRPO))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:implements ai:FSDP2ZeROSharding))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:implements ai:Int8AllReduce))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:implements ai:VerifiableInference))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:implements ai:AsynchronousRL))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:uses ai:LlamaArchitecture))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:uses ai:AkashNetwork))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:uses ai:RayDistributedComputing))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:uses ai:LocalitySensitiveHashing))Reduction Relationships
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:reduces ai:CommunicationBandwidthRequirements))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:reduces ai:TrainingCentralisation))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:reduces ai:ComputeAccessBarrier))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:reduces ai:AIMonopoly))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:reduces ai:SynchronisationOverhead))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:contrasts ai:CentralisedTraining))
SubClassOf(ai:PrimeIntellect
ObjectSomeValuesFrom(ai:contrasts ai:ProprietaryAI))
DataPropertyAssertion(ai:hasIdentifier ai:PrimeIntellect "AI-2047"^^xsd:string)
DataPropertyAssertion(ai:authorityScore ai:PrimeIntellect "0.87"^^xsd:decimal)
DataPropertyAssertion(ai:communicationReduction ai:PrimeIntellect "400"^^xsd:integer)
DataPropertyAssertion(ai:computeUtilisationPercent ai:PrimeIntellect "83"^^xsd:integer)
DataPropertyAssertion(ai:largestModelParameters ai:PrimeIntellect "106000000000"^^xsd:integer)
DataPropertyAssertion(ai:totalFundingUSD ai:PrimeIntellect "70400000"^^xsd:integer)
AnnotationAssertion(rdfs:label ai:PrimeIntellect "Prime Intellect"@en)
AnnotationAssertion(rdfs:comment ai:PrimeIntellect "Decentralised AI research organisation and GPU compute marketplace demonstrating frontier-scale distributed training via PRIME framework (DiLoCo+FSDP2 hybrid, 400x communication reduction, ElasticDeviceMesh), INTELLECT-1 (10B globally distributed LLM, Nov 2024), INTELLECT-2 (32B distributed RL, Apr 2025), INTELLECT-3 (106B MoE, Nov 2025), and METAGENE-1 (7B metagenomic pandemic model, Jan 2025), with $70.4M funding and open-source philosophy opposing centralised AI monopoly."@en)
AsymmetricObjectProperty(ai:enables)
AsymmetricObjectProperty(ai:requires)
TransitiveObjectProperty(ai:dependsOn)About Prime Intellect
- Prime Intellect is an open-source AI company and decentralised compute organisation founded in San Francisco in 2023, occupying a distinctive position at the intersection of distributed systems engineering, frontier model training, and the decentralised AI movement.
- Core mission: demonstrate that frontier-scale (10B–100B+ parameter) LLM training can be executed across geographically distributed, independently owned compute nodes contributed by parties worldwide, without a central GPU supercomputer or single controlling entity.
- Founders: Vincent Weisser (CEO) previously co-founded VitaDAO (leading decentralised longevity research organisation) and served as AI/product lead at Molecule (biopharma IP tokenisation platform)—bringing a decentralised science (DeSci) philosophy of open, community-governed scientific infrastructure. Johannes Hagemann (CTO) built distributed training frameworks at Aleph Alpha (German frontier AI lab) and received NeurIPS 2023 Best Paper for distributed optimisation work, providing technical credibility for attacking large-scale decentralised training.
- What makes Prime Intellect different from federated learning research: Earlier FL research (McMahan et al. FedAvg, Google Gboard) targeted privacy-preserving cross-device learning on mobile phones with tiny local datasets and simple models. Prime Intellect targets frontier-scale pre-training and post-training RL where models have 10B–106B+ parameters, training runs last weeks to months, and gradient communication is measured in terabytes per synchronisation round.
- Engineering challenges Prime Intellect addresses that prior FL research did not: (1) minimising cross-continental gradient synchronisation overhead by four orders of magnitude; (2) tolerating arbitrary node failures mid-run without restarting from epoch zero; (3) handling heterogeneous GPU architectures (H100, A100, RTX 3090, RTX 4090) and network speeds (10 Mbps to 100 Gbps) in a single training job; (4) verifying that remote compute contributors execute the correct computation honestly, detecting fraud or misconfiguration.
- Commercial structure: Prime Intellect operates simultaneously as an open-source AI research organisation (models, training code, and datasets under Apache 2.0) and a for-profit GPU compute marketplace. Revenue from the compute exchange cross-subsidises research training runs. Total $70.4M in funding has been allocated to engineering team growth, training run infrastructure, and compute acquisition.
- Open-source philosophy: All INTELLECT model weights, intermediate checkpoints, pre-training data, post-training data, and training code are released publicly. PRIME framework and PRIME-RL are open-source on GitHub. This distinguishes Prime Intellect from centralised frontier labs (OpenAI, Anthropic, Google DeepMind) that do not release pre-training code or data.
PRIME Framework: Technical Architecture
- PRIME (Prime Resilient Intelligent Mesh Engine) is the production distributed training framework underpinning all INTELLECT runs from INTELLECT-1 onwards, addressing three core engineering challenges of internet-scale distributed training: communication efficiency, fault tolerance, and heterogeneous node management.
Two-Level Optimisation Hierarchy
- Hierarchical design: PRIME separates fast intra-node communication from slow inter-node synchronisation using a two-level optimiser hierarchy inspired by Google DeepMind’s DiLoCo (Douillard et al. 2023).
- Inner optimiser: Standard AdamW SGD running within each compute node’s 8×H100 GPU island connected by NVLink/NVSwitch (900 GB/s NVSwitch bidirectional bandwidth for DGX H100). Within each node, PyTorch FSDP2 (Fully Sharded Data Parallel, second generation) with ZeRO-3 parameter sharding distributes model weights, gradients, and optimiser states across 8 GPUs.
- Memory mathematics: A 10B parameter model (fp16 weights ~20 GB; fp32 optimiser states ~80 GB) fits within 80 GB HBM2e per GPU × 8 = 640 GB effective per node with ZeRO-3 full sharding. A 32B model requires ~256 GB optimiser states, fitting on 4 nodes of 8×H100.
- Inner step accumulation: The inner optimiser runs for H=100 training steps (~40 minutes at 10B scale) entirely within the fast NVLink fabric, never touching the internet connection during these steps.
- Outer optimiser: After H inner steps, PRIME computes pseudo-gradients Δθ = θ_initial − θ_after_H_steps (the parameter difference before and after H local steps) and performs an all-reduce across all distributed nodes over the internet.
- Outer gradient aggregation: The outer synchroniser applies Nesterov momentum SGD to the aggregated pseudo-gradients: θ_global ← θ_global + α × outer_grad + β × momentum_term, following the DiLoCo formulation.
- Communication reduction: Because pseudo-gradients are transmitted once every 100 steps rather than every forward-backward pass, algorithmic communication volume is reduced by ~100× before any further compression. Combined with int8 quantisation (4× additional reduction), total reduction is ~400× versus standard fp32 data-parallel all-reduce.
Int8 All-Reduce Compression
- Problem: Standard fp32 all-reduce of a 10B parameter model transmits 40 GB per synchronisation round over the internet—impractical on inter-continental links (4–10 Gbps).
- Int8 quantisation scheme: Each pseudo-gradient tensor is quantised from fp32 to int8 using per-tensor scale factors s = max(|Δθ|) / 127. Quantised value: int8_val = round(fp32_val / s). This reduces 40 GB to 10 GB (4× compression).
- Bias correction: Int8 truncation systematically underestimates large gradient values (clipping at ±127/s). PRIME applies a bias-correction term after de-quantisation to recover unbiased gradient estimates, preserving fp32-equivalent convergence at 10B scale with no measurable perplexity degradation in ablation experiments.
- Custom CUDA kernel: The int8 quantisation kernel was custom-implemented achieving 60× speed improvement over naive Python int8 conversion—critical for keeping outer synchronisation under 1 minute.
- Extreme bandwidth reduction: In bandwidth-constrained experimental configurations using int8 all-reduce at full 100-step outer interval, effective bandwidth reduction versus standard fp32 data-parallel all-reduce reaches up to 2,000×.
ElasticDeviceMesh: Fault Tolerance
- Problem with standard PyTorch distributed: Standard
torch.distributedinitialises a fixed process group; any node failure causes an immediate collective communication timeout and full training abort, requiring restart from the last checkpoint. - ElasticDeviceMesh design: Maintains two parallel process group registries: (1) the global process group managing DiLoCo inter-node all-reduce over TCP/IP across the internet; (2) local process groups managing FSDP2 intra-node NVLink/InfiniBand communication. The two levels are independently fault-tolerant.
- Failure recovery sequence: (a) heartbeat timeout detection (configurable, default 60 seconds); (b) remove failed node from global process group; (c) transparently reshard model across remaining nodes using FSDP2 reshard; (d) broadcast latest valid checkpoint to any replacement node joining; (e) resume training without full restart.
- Critical for INTELLECT-1: 30 independent compute sponsors with varying reliability contributed nodes throughout the run. Individual contributor nodes dropped and rejoined repeatedly—ElasticDeviceMesh allowed the run to proceed continuously without requiring human intervention for each node failure.
- Dynamic node onboarding: New nodes can join a running training run mid-flight. ElasticDeviceMesh assigns the new node a shard of the FSDP2 model, streams the current checkpoint, and incorporates it into the next outer synchronisation without interrupting existing nodes.
Asynchronous Distributed Checkpointing
- Problem: Synchronous checkpointing of a 10B+ model blocks GPU training for 20+ minutes per checkpoint (serialising ~100 GB of fp32 model state + optimiser states across 8 GPUs over NFS or object storage).
- Async solution: Immediately after each outer synchronisation step, model state is atomically copied to CPU RAM on a background thread (NVLink bandwidth ~50 GB/s; copy completes in ~1 minute). GPU training resumes immediately on the next inner steps while the background thread writes CPU-RAM copy to persistent storage (NFS, S3, or distributed filesystem).
- Result: Checkpoint blocking time reduced from 20+ minutes to under 1 minute, recovering approximately 2–5% effective GPU utilisation for large runs. Checkpoints are verified via content hash before being marked valid—preventing corrupted checkpoints from being used in recovery.
DiLoCo and OpenDiLoCo
- DiLoCo (Distributed Low-Communication Training) is the Google DeepMind algorithm (Arthur Douillard et al., 2023) underlying PRIME’s outer optimiser. DiLoCo applies federated learning-inspired local optimisation to LLM pre-training, exploiting the empirical finding that LLM loss surfaces are smooth enough that local gradient updates over H=500 steps remain within the correct descent basin.
- DiLoCo key result: 8 workers with DiLoCo achieve parity with fully synchronous data-parallel training while reducing communication by 500×. The outer optimiser uses Nesterov momentum SGD applied to pseudo-gradients Δθ_i = θ^(t)_i − θ^(t-H)_global, where θ^(t)_i is worker i’s parameters after H local steps and θ^(t-H)_global is the last globally synchronised checkpoint.
- Why DiLoCo works: LLM gradients have a low-rank structure—the most informative gradient directions for pre-training change slowly over hundreds of steps, meaning infrequent global synchronisation captures essentially the same signal as step-by-step synchronisation at far lower communication cost.
- OpenDiLoCo (arXiv:2407.07852, July 2024): Prime Intellect’s open-source reproduction and extension of DiLoCo, released July 2024. Built on the Hivemind library (distributed hash table over libp2p for peer discovery without a central coordinator).
- OpenDiLoCo contributions vs DeepMind paper: (1) 3× parameter count scaling—first publicly reproducible evidence that DiLoCo generalises to billion-parameter models beyond DeepMind’s original experiments; (2) real-world cross-continental deployment across two continents and three countries—first empirical proof of DiLoCo over genuine internet connections with real-world latency and packet loss; (3) 90–95% compute utilisation in cross-continental experiments; (4) Apache 2.0 open-source code enabling community extension.
- OpenDiLoCo deprecation: OpenDiLoCo is now deprecated in favour of the production PRIME framework (which adds ElasticDeviceMesh fault tolerance, int8 all-reduce CUDA kernels, and full FSDP2 integration). The GitHub repository (PrimeIntellect-ai/OpenDiloco) directs users to prime-rl for production use.
- Decoupled DiLoCo (Google DeepMind, arXiv:2604.21428, April 2025): Subsequent extension decoupling outer gradient computation from node synchronisation barriers, improving resilience to node failures and providing additional theoretical validation of the DiLoCo paradigm.
INTELLECT-1: World’s First Globally Distributed 10B Foundation Model
- Announcement: October 11, 2024. Release: November 29, 2024. Technical report: arXiv:2412.01152, December 2024.
- Model architecture: 10B-parameter Llama-3 architecture language model. Llama-3 uses grouped-query attention (GQA) with 8 KV heads and 32 query heads, rotary position embeddings (RoPE), SwiGLU activation in FFN layers, and no bias terms—a lean, efficient architecture well-suited for distributed training where synchronisation overhead is critical.
- Training data mixture: 1 trillion tokens from four sources: 55% FineWeb-Edu (high-quality web text filtered for educational content by GPT-4-trained classifier), 20% DLCM (diverse language corpus mixture for multilingual coverage), 20% Stack v2 (code from 600+ programming languages, deduplicated and PII-redacted), 5% OpenWebMath (mathematical exposition and problem-solving content). Mixture optimised via 1B-scale ablation experiments.
- Training infrastructure: 30 independent compute sponsors contributing 8×H100 nodes—named sponsors include Hugging Face, SemiAnalysis, Arcee AI, Hyperbolic, Olas, Akash Network, Schelling AI, and multiple anonymous individual GPU contributors—spanning 5 countries on 3 continents (North America, Europe, Asia-Pacific).
- Training configuration details: DiLoCo outer sync interval H=100 inner steps; inner batch size 2,048 sequences × 2,048 tokens = ~4M tokens per inner step; outer sync all-reduce with int8 quantisation; WSD learning rate schedule (warm-up 2,000 steps, stable phase, cool-down 10,000 steps on higher-quality data subset).
- WSD learning rate schedule: Warmup-Stable-Decay scheduler with linear warm-up to peak learning rate 3×10^-4, stable plateau for the majority of training, then cosine decay to 10% peak rate over the final ~10% of tokens. The decay phase switches to a higher-quality filtered data subset (removing lower-quality web text), following findings from Mistral/Gemma that data quality elevation at training end improves benchmark performance disproportionately.
- Performance metrics: 83% overall compute utilisation across all contributors and geographies; 96% compute utilisation within the United States (inter-data-centre bandwidth ~4 Gbps); all-reduce synchronisation completing in under 1 minute per outer step; communication representing only 1–2% of total training time.
- Communication efficiency breakdown: Baseline fp32 all-reduce: 40 GB per 10B model → 400 MB after 100-step outer interval → 100 MB after int8 quantisation = 400× total reduction. US-region all-reduce time at 4 Gbps: 100 MB / 500 MB/s = 0.2 seconds per sync. Transcontinental (4 Gbps → 1 Gbps): 100 MB / 125 MB/s = 0.8 seconds per sync. Both well within the <1 minute target.
- Open release scope: Base model checkpoint, instruct fine-tune (RLHF-lite with 5K human preference annotations), all 50+ intermediate checkpoints (every outer sync saved), pre-training data recipe (sampling weights, source URLs), post-training data (instruction pairs), and complete training code including ElasticDeviceMesh implementation—all Apache 2.0 on Hugging Face.
- INTELLECT-1 significance: First empirical proof that a frontier-scale (10B+) LLM can be collaboratively trained by geographically distributed, independently owned compute without a central supercomputer. Establishes practical upper bound on DiLoCo communication efficiency (400× reduction) and lower bound on distributed compute utilisation (83% globally, 96% regionally).
- Technical report ablation studies: H comparison (H=10 vs H=50 vs H=100 vs synchronous): H=100 adds <0.5 perplexity points vs synchronous on standard LM benchmarks. Int8 vs fp32 outer sync: no measurable perplexity difference at 10B scale. Node count scaling (4 vs 8 vs 16 nodes): linear throughput scaling with <5% overhead per doubling of node count. Fault recovery overhead: <2% additional compute time from ElasticDeviceMesh recovery operations across the full training run.
INTELLECT-2: First Globally Distributed RL Training of a 32B Reasoning Model
- Announcement and release: April 15, 2025. Technical report: arXiv:2505.07291, May 2025.
- Model: QwQ-32B (32B parameters, Qwen-family architecture) fine-tuned via globally distributed asynchronous GRPO reinforcement learning.
- Core innovation: First globally distributed RL training run at 32B scale with permissionless heterogeneous compute contributions—participants can contribute with as little as 4×RTX 3090 consumer GPUs, not H100s.
- GRPO (Group Relative Policy Optimization): The RL algorithm from DeepSeek-R1 (January 2025). For each prompt x, G=8 completions are sampled from the current policy; rewards r_i are computed from verifiers (maths checkers, code executors); advantages are group-normalised A_i = (r_i − mean(r)) / std(r); policy gradient: clip(π(completion_i|x)/π_ref, 1−ε, 1+ε) × A_i. GRPO eliminates the critic/value network that would double 32B-scale memory requirements—making it tractable for distributed consumer hardware.
- Asynchronous RL design: Inference workers generate rollouts from policy versions up to 4 gradient steps behind the current training checkpoint. Experiments demonstrate policy lag of up to 4 steps induces no measurable performance degradation, enabling variable-latency node participation without blocking training on the slowest worker.
- Three-tier architecture: (1) Inference rollout workers—4×RTX 3090 minimum, generating completions; (2) TOPLOC validators—verifying rollout authenticity; (3) GRPO training workers—centralised training nodes computing advantages and updating policy.
- Training data and online filtering: Difficulty-filtered dataset (≤75% solve rate problems only—eliminating trivial examples where all rollouts get identical rewards, contributing zero gradient signal); online advantage filtering removes batches where advantage variance falls below threshold, detecting reward hacking.
- Controllable thinking: Length rewards teach model to adhere to token budgets specified in system prompts, replicating the L1 paper (Aggarwal et al.) controllable-reasoning finding in a distributed RL setting.
- Performance: Improved mathematics and coding benchmark performance over base QwQ-32B; DeepScaleR reasoning benchmark replication without model degradation; predictable reasoning performance scaling with token budget.
TOPLOC: Verifiable Inference
- Problem: In distributed RL, remote inference workers must generate honest rollouts. A misconfigured or malicious node could submit rollouts from a different model (lower quality, lower precision, different prompt), corrupting the training gradient signal.
- Solution: TOPLOC (TOPology LOCality hashing) computes a compact hash signature over L2-normalised top-K intermediate layer activations during each forward pass.
- How it works: (1) During rollout generation, worker computes TOPLOC hash over selected transformer layer activations; (2) hash is transmitted alongside the rollout to the central validator; (3) validator recomputes hash over submitted rollout’s intermediate activations and checks for match.
- Key properties: (a) Robust to GPU non-determinism—small floating-point differences between RTX 3090 and H100 (different CUDA kernel rounding) are within the locality-sensitive hash tolerance; (b) Sensitive to meaningful deviations—model weight modification, prompt substitution, or precision change (fp16→fp8) detected with high probability; (c) Faster than generation—validator runs a single deterministic forward pass, not full autoregressive sampling; (d) No trusted hardware required—purely software-based verification without TEE (trusted execution environment) or SNARK computation.
- Mathematical basis: Locality-sensitive hashing (Indyk & Motwani 1998)—hash family where similar activation vectors have high collision probability and dissimilar vectors have low collision probability, enabling anomaly detection via hash mismatch.
SHARDCAST: Distributed Weight Broadcast
- Problem: After each GRPO update, all 32B policy weight parameters (fp16: ~64 GB) must be pushed to all inference rollout workers before next rollout generation round. Naïve unicast from training origin to N=50+ workers requires 64 GB × N transmissions from the training nodes—infeasible.
- Solution: HTTP-based tree-topology broadcast network. Three tiers: (1) origin server (master weight copy after training step); (2) middle nodes (6–10 high-bandwidth relay nodes, cache and rebroadcast to worker subsets); (3) client nodes (rollout workers, pull from nearest middle node).
- Result: O(log N) reduction in origin egress bandwidth; 32B policy weights propagated to 50+ workers in under 5 minutes even with large models. Middle nodes are selected by Prime Intellect from high-bandwidth, high-uptime contributors to the compute network.
PRIME-RL: Agentic RL Training Framework
- PRIME-RL is Prime Intellect’s open-source agentic RL framework (GitHub: PrimeIntellect-ai/prime-rl), released alongside INTELLECT-3 in November 2025.
- Design goal: Scale seamlessly from single GPU to 1,000+ GPUs with first-class support for agentic multi-turn environments where the model interacts with external tools across extended episodes before receiving a terminal reward signal.
- Training backend: PyTorch FSDP2 for model weight sharding; FP8 inference via transformer-engine (2× throughput vs fp16 on H100/H200); expert parallelism (EP) and context parallelism (CP) for MoE models; PD disaggregation (prefill-decode separation for efficient variable-length RL rollout batching).
- Expert parallelism (EP): In MoE models, different experts are assigned to different GPU groups. EP partitions expert weight matrices across GPU nodes, allowing the MoE routing mechanism to dispatch tokens to their designated expert’s GPU without replicating all expert weights on all GPUs. Critical for 106B MoE where full expert replication would exceed per-node memory.
- Context parallelism (CP): Partitions long-context sequences across GPU groups along the sequence dimension, enabling training on very long context windows (up to 128K tokens) that would not fit in a single GPU’s KV-cache. Essential for agentic tasks where multi-turn tool-use episodes accumulate large context.
- Inference backend: vLLM with PagedAttention KV-cache management for 10–50× higher rollout throughput vs naive autoregressive sampling. PagedAttention manages KV-cache memory in non-contiguous pages, eliminating the memory fragmentation that makes naive batched inference inefficient for variable-length sequences.
- Verifiers and Environments Hub: Unified environment interface supporting maths verifiers (SymPy symbolic equality checking for exact mathematical answers), code execution (isolated sandboxes for Python/C++/Rust unit test evaluation), science benchmarks (multi-choice science reasoning), and custom agent environments. Community-extensible ecosystem allowing researchers to contribute new RL task definitions.
- Prime Sandboxes: High-throughput isolated code execution environments enabling >10,000 parallel code evaluations/second across a distributed cluster—critical for RL training where rewards come from unit test pass rates on model-generated code. Isolation prevents malicious or buggy model-generated code from affecting the training infrastructure.
- Key design property: Framework generalises from distributed internet training (INTELLECT-2 asynchronous GRPO) to centralised high-performance clusters (INTELLECT-3 512×H200 RL), using the same codebase without architectural changes—a key software engineering achievement that reduces maintenance burden across Prime Intellect’s training portfolio.
INTELLECT-3: 106B MoE with Large-Scale RL
- Release: November 2025. Technical report: arXiv:2512.16144. Base model: GLM-4.5-Air (106B MoE from Tsinghua/Zhipu AI).
- Training stages: (1) Supervised fine-tuning (SFT) on high-quality maths, code, and reasoning demonstrations; (2) large-scale RL stage using GRPO with PRIME-RL on 512×H200 GPU cluster (64 nodes) over approximately two months.
- Infrastructure: 512 NVIDIA H200 GPUs, each with 141 GB HBM3 (vs 80 GB HBM2e for H100), enabling larger batch sizes and more aggressive sharding. FSDP2 with FP8 inference and expert parallelism for MoE routing.
- Benchmark performance: State-of-the-art for 106B MoE size class on AIME 2024 mathematics, LiveCodeBench coding, and GPQA Diamond science reasoning benchmarks.
- Open-source scope: Model weights, full training recipe, PRIME-RL framework, verifiers, and Environments Hub released under Apache 2.0.
- Centralised vs decentralised note: INTELLECT-3 was trained on a centralised cluster (not over the internet), a pragmatic acknowledgement that 106B-scale RL required coordination levels beyond pure DiLoCo in 2025. Prime Intellect positions INTELLECT-3’s contribution as the open-source RL tooling rather than the training topology, with decentralised 100B+ RL targeting INTELLECT-4.
METAGENE-1: Open Science Foundation Model for Pandemic Monitoring
- Release: January 2025. Technical report: arXiv:2501.02045. Collaboration: Prime Intellect + University of Southern California (USC) + Nucleic Acid Observatory.
- Architecture: 7B-parameter autoregressive transformer (GPT-2-style at large scale), nucleotide-level tokenisation (A/T/G/C/N tokens with k-mer context windows).
- Training corpus: 1.5 trillion DNA and RNA base pairs from large-scale human wastewater metagenomic sequencing—thousands of samples across diverse geographic locations processed via Illumina deep metagenomic sequencing and quality filtration. This is environmental sequencing data, not reference genome databases.
- Why wastewater metagenomics: Environmental metagenomics captures the full microbial diversity present in human populations—including novel pathogens, emerging variants, and unculturable organisms invisible to reference databases. Wastewater surveillance detects population-level disease signals before clinical cases are reported.
- Pathogen detection performance: 92.96 Matthews Correlation Coefficient (MCC) on wastewater pathogen detection benchmark—significantly outperforming previous state-of-the-art (~78 MCC). MCC preferred over accuracy due to severe class imbalance (pathogens are rare vs commensal microbes).
- Additional benchmarks: State-of-the-art on metagenomic embedding tasks, phylogenetic classification, and antimicrobial resistance gene detection.
- Dual-use safety: Architecture prevents harmful misuse—512-token context window and no generative fine-tuning on pathogen genomes means METAGENE-1 cannot be used to design novel pathogens. Founders call this “differential technological development”: defence-favouring AI applications that benefit biosurveillance without enabling bioweapons design.
- Open release: Model weights, training data processing pipeline, and inference code released publicly.
GPU Compute Exchange
- Concept: Prime Intellect’s compute marketplace functions as a meta-cloud aggregating GPU supply from multiple source categories into a unified programmatic marketplace with transparent pricing—a compute exchange analogous to financial exchanges that aggregate supply and demand across market participants.
- Market inefficiency addressed: The GPU rental market in 2024–2026 exhibits significant price discrepancies (identical H100 GPU hours vary by 40–70% across providers), low price transparency (most providers require sales conversations rather than publishing live rates), and fragmented supply (hundreds of independent providers with no unified API). Prime Intellect’s exchange directly addresses these inefficiencies.
- Supply tier 1—Public cloud reselling: AWS, GCP, Azure, OCI GPU instances aggregated with value-added orchestration. Prime Intellect provides SLURM/Kubernetes scheduling, distributed training tooling, and observability on top of raw cloud GPU instances, justifying a margin over bare cloud rates.
- Supply tier 2—Independent data centres: Colocation facilities and smaller GPU cloud providers (Hyperbolic, CoreWeave, Lambda Labs, Massed Compute) list available inventory on the Prime Intellect marketplace API, with Prime Intellect handling customer-facing scheduling and support.
- Supply tier 3—Akash Network integration: Prime Intellect integrates Akash Network’s permissionless blockchain-based GPU marketplace, where GPU providers register on-chain (Cosmos-based blockchain) and receive AKT token payments. Prime Intellect translates between Akash’s smart-contract-mediated bidding system and the standard REST API that ML researchers use to provision compute, handling AKT/USD conversion and regulatory compliance.
- Pricing mechanisms: Live bid/ask order book by GPU type (H100 SXM5, H100 PCIe, A100 80GB, RTX 4090, RTX 3090), GPU count, interconnect type (InfiniBand vs Ethernet vs NVLink topology), and geographic region. Spot/preemptible options (cheapest, subject to preemption), on-demand (fixed rate, no preemption), and reserved (discounted long-term contracts, 1-month to 12-month).
- Provider reliability metrics: Published openly: uptime percentage (30-day rolling), job completion rate, median job start latency, VRAM tested stable, network bandwidth measured. Buyers can explicitly trade off cost against reliability—mission-critical training runs filter for >99% uptime providers; research exploration workloads can tolerate lower reliability at lower cost.
- Workload orchestration: SLURM and Kubernetes job scheduling; custom Docker image support for reproducible training environments; InfiniBand networking for multi-node low-latency jobs; Ray distributed computing integration for complex multi-process workloads; Grafana real-time metrics dashboards for GPU utilisation, memory pressure, network throughput, thermal throttling, and job health.
- Scale supported: 1–256+ GPU workloads; designed for pre-training, post-training (SFT, RL), and agent evaluation—not merely inference. Large runs can request multi-day or multi-week cluster reservations.
- Prime Intellect API: Programmatic REST API with Python SDK for automated resource acquisition, job submission, monitoring, status callbacks, and teardown. Enables research CI/CD pipelines that programmatically spin up and tear down GPU clusters—a capability absent from traditional HPC procurement processes.
Use Cases and Major Applications
- Collaborative frontier pre-training: INTELLECT-1 proves a 10B LLM trained on 1T tokens can be produced without any entity owning a supercomputer—enabling research groups, universities, and DAOs to contribute compute to frontier model training runs they otherwise could not afford.
- Permissionless RL at 32B scale: INTELLECT-2’s minimum of 4×RTX 3090 consumer GPUs for RL rollout generation democratises post-training reinforcement learning participation—previously accessible only to organisations with large H100 clusters.
- GPU marketplace as distributed infrastructure: Independent GPU owners (individual miners, small data centres, university clusters) list spare capacity; Prime Intellect aggregates supply into a more competitive GPU rental market, addressing the fragmentation and price opacity of the current GPU cloud market.
- Open science foundation models: METAGENE-1 demonstrates that frontier-scale (7B) domain-specific models too expensive for most academic groups to train alone can be produced through open collaborative infrastructure and released for high-impact public health applications.
- Decentralised AI governance research: Prime Intellect’s tokenised ownership protocol roadmap (2026–2027) proposes a model for cooperative AI development where compute contributors receive fractional ownership stakes, enabling non-commercial entities and individual researchers to participate in frontier model governance.
Academic Context
- DiLoCo (Douillard et al., DeepMind, 2023): Direct technical ancestor of PRIME’s outer optimiser. Established that Nesterov outer SGD on pseudo-gradients achieves 500× communication reduction while maintaining LLM pre-training convergence. Prime Intellect extended DiLoCo with int8 quantisation (4× additional compression), FSDP2 intra-node sharding, and ElasticDeviceMesh fault tolerance.
- Federated Averaging (McMahan et al., AISTATS 2017): Conceptual ancestor establishing that local optimisation over H steps before global aggregation is compatible with learning—DiLoCo applies this insight at frontier training scale. FedAvg targeted mobile devices with millions of participants; DiLoCo/PRIME targets tens of GPU nodes with billions of parameters per node.
- PyTorch FSDP2 (Zhao et al., arXiv:2304.11277, 2023): Intra-node memory efficiency backbone of PRIME’s inner optimiser. DTensor-based parameter partitioning supports dynamic resharding when node count changes—critical for ElasticDeviceMesh’s fault-tolerant reshard on node failure.
- ZeRO-3 (Rajbhandari et al., SC2020): Parameter, gradient, and optimiser state partitioning across GPUs within a node. Three-stage memory reduction: ZeRO-1 (optimiser states), ZeRO-2 (+ gradients), ZeRO-3 (+ parameters). PRIME uses ZeRO-3 (FSDP2) enabling 10B-scale models on 8×H100 without pipeline or tensor parallelism complexity.
- GRPO / DeepSeek-R1 (DeepSeek-AI, January 2025): The RL algorithm adopted by INTELLECT-2. Key properties making GRPO tractable for distributed RL: no critic network (halves memory vs PPO); group normalisation eliminates reward scale sensitivity; online advantage filtering handles asynchronous rollout batches.
- Hivemind (Ryabinin et al., NeurIPS 2020): DHT-based peer discovery library used in OpenDiLoCo. PRIME replaced Hivemind’s DHT with a custom TCP-based global process group: deterministic latency (DHT lookup adds variable latency), simpler failure semantics (DHT entry expiry vs explicit heartbeat), and tighter integration with FSDP2’s process group model.
- Locality-Sensitive Hashing (Indyk & Motwani, STOC 1998): Mathematical foundation for TOPLOC’s verifiable inference. Random projection LSH maps high-dimensional activation vectors to low-dimensional signatures: Pr[h(x) = h(y)] = 1 − (1/π) × arccos(cos(θ_{xy})) for unit-normalised vectors x, y with angle θ_{xy}. Similar activations (from the same model on the same prompt) have high hash collision probability; dissimilar activations (from a different model or modified prompt) have low collision probability.
- Mixture-of-Experts (Shazeer et al., arXiv:1701.06538, 2017): Architecture foundation for INTELLECT-3’s 106B MoE. Sparse MoE routes each token to a subset of expert FFN layers (top-2 routing), activating only ~20B parameters per forward pass despite 106B total parameters—enabling 106B-capacity models at the compute cost of a ~20B dense model.
- Controllable Reasoning (Aggarwal et al., L1 paper, 2025): Length reward technique replicated in INTELLECT-2’s controllable thinking: reward model assigns additional reward for completing within a token budget specified in the system prompt, training the model to allocate thinking tokens proportionally to problem difficulty.
Current Landscape (2026)
- Prime Intellect’s trajectory: OpenDiLoCo (July 2024) → INTELLECT-1 10B (November 2024) → INTELLECT-2 32B RL (April 2025) → INTELLECT-3 106B MoE RL (November 2025): 18 months from proof-of-concept to 100B+ scale, all open-source.
- Competitive landscape in decentralised AI training: Petals (BigScience/Yandex)—distributed inference over volunteer networks, targets inference not pre-training; Bittensor (TAO)—blockchain compute network with on-chain consensus, latency incompatible with gradient-synchronous training; Together AI/CoreWeave—aggregated GPU cloud without decentralised ownership; Nous Research—open decentralised fine-tuning experiments.
- Prime Intellect differentiators: Targets pre-training and RL (not inference); TOPLOC-verified honest computation; consistent Apache 2.0 open-sourcing of models and code; INTELLECT series as empirical proof points at increasing scale.
- Centralised training still dominates: The majority of frontier model training in 2026 occurs on centralised GPU clusters—NVIDIA DGX SuperPOD configurations (4,000–10,000 H100s), national AI compute programmes (UK AIRRC, French GENCI, US NAIRR), and hyperscaler proprietary clusters. Decentralised approaches remain a growing but small share of total frontier compute.
- Tension in the thesis: INTELLECT-3 was trained on a 512×H200 centralised cluster—not over the internet—a pragmatic acknowledgement that 106B-scale RL required more coordination than pure internet-distributed DiLoCo could provide in 2025. Critics note the gap between the decentralisation philosophy and the centralised reality of the flagship training run. Prime Intellect’s response: INTELLECT-3’s contribution is the open-source RL tooling, not the topology; INTELLECT-4 will target fully decentralised RL at 100B+ scale.
- GPU market context: Prime Intellect’s marketplace operates in a GPU rental market characterised by significant price discrepancies between providers, low price transparency, and fragmented supply. The marketplace addresses these structural inefficiencies by aggregating supply and publishing live pricing—a role analogous to cloud broker services but targeting ML training workloads specifically.
UK Context
- UCL Centre for AI and Gatsby Unit: UCL’s Centre for AI (DeepMind co-funded) and Gatsby Computational Neuroscience Unit maintain research programmes in federated learning and communication-compressed distributed optimisation directly relevant to DiLoCo/PRIME. UCL co-leads the Generative AI Hub consortium (UCL, Imperial, Cambridge, Oxford, Manchester, Edinburgh, Cardiff, Surrey) providing shared UK academic GPU compute—a national academic analogue to Prime Intellect’s marketplace model at smaller scale. UCL’s MSc Machine Learning, one of Europe’s longest-running specialist ML programmes, produces graduates with strong distributed systems foundations for decentralised training research.
- Imperial College London Machine Learning Initiative: Imperial’s ML Initiative and the Security & Machine Learning Lab operate a hybrid GPU infrastructure combining community-run NVIDIA RTX 40/50-series GPUs with high-end A100/H200 resources and AWS Trainium/Inferentia nodes—heterogeneous compute topology structurally analogous to Prime Intellect’s marketplace. Imperial has piloted integration of Theta EdgeCloud decentralised compute into research workloads, exploring the same permissionless compute model Prime Intellect operationalises at frontier scale. The Intelligent Transmission and Processing (ITP) Lab conducts research in distributed optimisation and communication-efficient methods directly relevant to DiLoCo’s outer synchronisation protocol.
- Edinburgh EPCC: The Edinburgh Parallel Computing Centre operates the national ARCHER2 HPC system (28-petaflop Cray EX supercomputer), providing expertise in large-scale distributed training across slow interconnects structurally analogous to Prime Intellect’s internet-distributed training challenge, though on dedicated academic networks (JANET/GÉANT ~100 Gbps) rather than commodity internet (4 Gbps). EPCC’s checkpoint management, fault tolerance, and heterogeneous accelerator orchestration experience maps directly to PRIME’s engineering challenges.
- Cambridge and ARM Holdings: ARM’s Cambridge R&D centres are developing next-generation AI accelerators with lower power consumption per FLOP than x86/NVIDIA combinations—attractive for Prime Intellect’s aspiration to aggregate always-on consumer hardware (gaming PCs, edge devices). ARM-based GPU alternatives (Apple M-series, Qualcomm AI 100, NVIDIA Jetson) are increasingly competitive for inference and small-scale fine-tuning, expanding the accessible hardware base for decentralised training in 2027–2030.
- Northern England industrial context: Manchester hosts Graphcore (IPU alternative AI accelerators) and the Northern AI Corridor initiative (Manchester–Leeds–Sheffield) constructing regional GPU cluster infrastructure—exactly the type of independent data centre supply Prime Intellect’s marketplace is designed to aggregate. Leeds Digital Health and Sheffield AMRC (Advanced Manufacturing Research Centre) represent sectoral applications where access to Prime Intellect’s open models and marketplace compute enables AI adoption without hyperscaler dependency. Newcastle’s Data City and the wider Northern English tech ecosystem are potential participant nodes in any future UK-focused INTELLECT training run.
Future Directions (2026–2030)
- INTELLECT-4 and fully decentralised 100B+ RL: Targeting 100B+ parameter models trained via genuinely internet-distributed RL—policy gradient updates distributed globally, not merely inference rollouts. Requires: improved DiLoCo outer optimiser for MoE routing consistency across nodes; higher-bandwidth outer coordination protocols; TOPLOC extensions for multi-turn agentic rollout verification; and advances in outer sync frequency reduction for high-latency links.
- End-to-end agentic training: PRIME-RL’s Environments Hub enables multi-turn tool-use RL where models interact with code interpreters, web browsers, file systems, and APIs across extended episodes. Scaling agentic training to decentralised contributions requires TOPLOC extensions for multi-turn trajectory verification and credit assignment methods correctly attributing long-horizon rewards to specific actions in distributed computation.
- Decentralised model merging via DiLoCo: Future runs will experiment with merging independently RL-fine-tuned model branches (maths specialist, code specialist, science specialist) via DiLoCo-style outer averaging—enabling asynchronous model development pipelines without synchronised multi-node training coordination. Each contributor can fine-tune specialised branches that are periodically merged into a shared foundation.
- Retail compute integration: Training on gaming PCs, consumer GPUs, and Starlink-connected edge nodes. Technical prerequisites: further outer sync interval extension (H > 100 for Starlink multi-second latency); TOPLOC-grade verification for consumer GPU non-determinism; economic models making small contributions viable after electricity costs. NVIDIA RTX 5090 (January 2025 launch, 3,352 TFLOPS fp16) makes consumer contributions increasingly meaningful at the outer DiLoCo sync level.
- Tokenised ownership protocol (2026–2027): Compute contributors receive fractional model ownership stakes tracked on-chain. Implementation challenges: fair contribution valuation (compute hours vs data curation vs code development); legal IP ownership treatment across jurisdictions; governance mechanisms for model update decisions. Inspiration: VitaDAO’s token-gated IP licensing model (co-founder Weisser’s prior work) suggests DAO-based governance with on-chain contribution tracking.
- Open science model expansion: Following METAGENE-1, Prime Intellect plans additional open science foundation models in domains where data is abundant but compute is scarce for academics: climate modelling, materials science, protein structure prediction, drug-target interaction, satellite imagery analysis.
Research and Literature
| Reference | Details |
|---|---|
| Douillard et al. (2023) | “DiLoCo: Distributed Low-Communication Training of Language Models.” Google DeepMind. Foundational outer-optimiser algorithm for PRIME. |
| Jaghouar et al. (2024a) | “OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication Training.” arXiv:2407.07852, August 2024. 3× DiLoCo scaling; cross-continental deployment; 90–95% utilisation. |
| Jaghouar et al. (2024b) | “INTELLECT-1 Technical Report.” arXiv:2412.01152, December 2024. 10B globally distributed LLM; 400× communication reduction; 83% compute utilisation. |
| Prime Intellect Team (2025a) | “INTELLECT-2: A Reasoning Model Trained Through Globally Decentralised Reinforcement Learning.” arXiv:2505.07291, May 2025. 32B distributed RL; TOPLOC; SHARDCAST; asynchronous GRPO. |
| Prime Intellect Team (2025b) | “INTELLECT-3: Technical Report.” arXiv:2512.16144, November 2025. 106B MoE; large-scale RL; PRIME-RL open-sourced. |
| Senghaas et al. (2025) | “METAGENE-1: Metagenomic Foundation Model for Pandemic Monitoring.” arXiv:2501.02045, January 2025. 7B on 1.5T DNA/RNA base pairs; 92.96 MCC. |
| Douillard et al. (2025) | “Decoupled DiLoCo: Resilient, Distributed AI Training at Scale.” arXiv:2604.21428, April 2025. Google DeepMind failure-resilient DiLoCo extension. |
| DeepSeek-AI (2025) | “DeepSeek-R1: Incentivising Reasoning Capability in LLMs via RL.” arXiv:2501.12948, January 2025. GRPO algorithm adopted by INTELLECT-2. |
| McMahan et al. (2017) | “Communication-Efficient Learning of Deep Networks from Decentralised Data.” (FedAvg). AISTATS 2017. Conceptual ancestor of DiLoCo’s local-step approach. |
| Zhao et al. (2023) | “PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.” arXiv:2304.11277. Intra-node memory efficiency backbone of PRIME. |
| Rajbhandari et al. (2020) | “ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.” SC2020. ZeRO-3 sharding in PRIME FSDP2. |
| Ryabinin et al. (2020) | “Towards Crowdsourced Training of Large Neural Networks Using Decentralised Mixture-of-Experts.” (Hivemind). NeurIPS 2020. DHT peer discovery in OpenDiLoCo. |
| Shazeer et al. (2017) | “Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.” arXiv:1701.06538. MoE architecture foundation for INTELLECT-3. |
| Aggarwal et al. (2025) | “L1: Controlling How Long a Reasoning Model Thinks with Reinforcement Learning.” Controllable thinking replicated in INTELLECT-2. |
| Indyk & Motwani (1998) | “Approximate Nearest Neighbours: Towards Removing the Curse of Dimensionality.” STOC 1998. Mathematical basis for TOPLOC locality-sensitive hashing. |
| Prime Intellect (2024a) | “INTELLECT-1: Launching the First Globally-Distributed Training of a 10B Parameter Model.” Blog post, October 2024. |
| Prime Intellect (2024b) | “OpenDiLoCo.” Blog post, July 2024. Open-source DiLoCo release. |
| Prime Intellect (2024c) | “Introducing Prime Intellect Compute: The Compute Exchange.” Blog post, 2024. |
| Prime Intellect (2025c) | “INTELLECT-2 Release.” Blog post, April 2025. |
| Prime Intellect (2025d) | “INTELLECT-3: A 100B+ MoE Trained with Large-Scale RL.” Blog post, November 2025. |
| Prime Intellect (2025e) | “METAGENE-1: Metagenomic Foundation Model.” Blog post, 2025. |
| CoinFund / Distributed Global (2024) | “Prime Intellect Secures $5.5M Seed Funding.” PR Newswire, April 23 2024. |
| Founders Fund (2025) | “Prime Intellect Series A, $15M.” Announced February 28, 2025. Peter Thiel’s Founders Fund, Menlo Ventures, and angel investors. |
| Akash Network (2025) | “Prime Intellect Integrates Permissionless Akash GPUs.” Akash Network blog. |
| Nebius (2025) | “Prime Intellect: Distributed Training and RL with NVIDIA GB200 NVL72.” Customer story. |
| Sacra (2025) | “Prime Intellect: Funding, news & analysis.” $70.4M total funding confirmed. |
| Cognitive Revolution Podcast (2024) | “Distributed Training, Decentralised AI: Prime Intellect’s Master Plan.” Interview with Weisser and Hagemann. |
| UCL Centre for AI / Generative AI Hub (2025) | UK academic AI compute consortium documentation. |
| Imperial College / Theta EdgeCloud (2025) | “Imperial College London to Use Theta EdgeCloud in its AI Research.” Medium/Theta Labs. |
Centralised vs Decentralised Training: Comparative Analysis
- Problem framing: Large language model pre-training at the 10B–100B scale requires communicating model gradients across a large number of GPUs. In centralised training, all GPUs reside in the same data centre connected by ultra-high-bandwidth fabric (InfiniBand HDR/NDR: 200–400 Gbps per link), allowing synchronous all-reduce of fp32 gradients after every forward-backward pass. In decentralised training, nodes are connected by the commodity internet (4–100 Gbps between data centres, 1–10 Gbps for consumer nodes), where synchronous fp32 all-reduce every step would consume 100% of network bandwidth on a 10B model.
- Bandwidth comparison: Standard fp32 data-parallel all-reduce for a 10B model transmits 40 GB per step. At 1 step/second (typical for 10B on 8×H100), this requires 320 Gbps bandwidth per node—far exceeding typical internet connectivity. DiLoCo reduces this to 40 GB per 100 steps = 400 MB/step effective average; int8 quantisation further reduces to 100 MB/step average. At 100 Mbps internet, 100 MB can be transmitted in 8 seconds—acceptable for 40-minute outer sync cycles.
- Compute utilisation comparison: A monolithic 8,000×H100 DGX SuperPOD achieves ~94–97% hardware utilisation (3–6% overhead from NVLink all-reduce, NCCL coordination, and I/O). INTELLECT-1 achieved 83% globally (17% overhead from internet latency, fault recovery, checkpointing) and 96% within a single US region—approaching centralised performance for intra-regional training. The remaining 4% gap from centralised training in the regional case is primarily fault-tolerance overhead.
- Cost comparison: A DGX H100 SuperPOD (1,024 H100s) costs approximately 2–5M/year to operate (power, cooling, facilities). Renting equivalent 1,024 H100 compute from cloud providers costs 24–48M/year). Prime Intellect’s marketplace prices are typically 20–40% below major cloud providers for comparable hardware, through aggregating independent supply with lower overhead. For a 10B model trained on 1T tokens requiring ~320M H100-hours, centralised cloud cost would be ~$3M; decentralised cost via Prime Intellect’s marketplace was significantly less through hardware diversity and spot pricing.
- Communication topology comparison: In centralised training, ring-allreduce or tree-allreduce over InfiniBand fabric achieves O(N) bandwidth scaling with N GPUs. In PRIME’s DiLoCo, communication is O(1) bandwidth-independent of the number of nodes—each outer sync transmits the same amount of data regardless of whether there are 4 or 100 nodes, because each node independently computes its local update and contributes a fixed-size pseudo-gradient. This O(1) scaling is a fundamental advantage of the DiLoCo approach for very large distributed systems.
- Model quality comparison: INTELLECT-1’s perplexity on standard benchmarks (Hellaswag, ARC-Challenge, MMLU, WinoGrande) is comparable to synchronously trained Llama-3 10B models on equivalent data—demonstrating that the communication reduction does not compromise model quality at this scale. The INTELLECT-1 technical report (arXiv:2412.01152) includes ablation experiments comparing DiLoCo H=100 outer sync with H=10, H=50, and fully synchronous (H=1) baselines, confirming that H=100 introduces less than 0.5 perplexity points of degradation versus synchronous training.
- Fault tolerance comparison: Centralised training on a dedicated cluster achieves near-zero node failure rate (0.01–0.1% hourly failure probability for enterprise H100s). Decentralised training across 30 independent sponsors experiences much higher failure rates (individual contributors may have 1–5% hourly failure probability, cloud spot instances may be preempted at 2–10% hourly rate). ElasticDeviceMesh handles these failures transparently; the question is whether recovery overhead exceeds the cost advantage of heterogeneous supply.
PRIME vs Competing Distributed Training Approaches
- PRIME vs Megatron-LM (NVIDIA): Megatron-LM implements 3D parallelism (tensor parallelism × pipeline parallelism × data parallelism) optimised for centralised DGX clusters with InfiniBand. Tensor parallelism requires ~400 Gbps all-to-all bandwidth per layer—completely infeasible over the internet. PRIME uses only data parallelism at the inter-node level (DiLoCo outer sync) and FSDP2 within nodes—eliminating tensor and pipeline parallelism from the inter-node communication, enabling internet-speed outer sync.
- PRIME vs DeepSpeed (Microsoft): DeepSpeed ZeRO stages 1–3 partition optimiser states, gradients, and parameters across GPUs for memory efficiency, but still require synchronous per-step all-reduce over the training fabric. DeepSpeed does not natively support infrequent outer synchronisation. PRIME’s FSDP2 (equivalent to ZeRO-3) operates within each node; the outer DiLoCo sync is orthogonal to and compatible with intra-node ZeRO sharding.
- PRIME vs Hivemind (original): Hivemind (Ryabinin et al. 2020) pioneered internet-scale distributed training using DHT-based peer discovery and delay-tolerant parameter averaging. Hivemind supports various averaging frequencies including asynchronous averaging (each peer syncs when convenient). PRIME builds on Hivemind’s foundational ideas but replaces the DHT coordination with a more deterministic TCP-based global process group and adds FSDP2 intra-node sharding, custom int8 CUDA kernels, and ElasticDeviceMesh fault tolerance unavailable in Hivemind.
- PRIME vs Petals: Petals (Borzunov et al. 2022) targets distributed inference of LLMs across volunteer nodes, routing different layers to different nodes (pipeline parallelism for inference). Petals cannot be directly applied to training because training pipeline parallelism requires backward-pass gradient communication with the same bandwidth requirements as tensor parallelism—too high for internet links. PRIME avoids this by using DiLoCo’s local optimisation approach which makes training gradient communication infrequent.
- PRIME vs Bittensor: Bittensor operates a blockchain-based compute network where miners process AI inference tasks and receive TAO token rewards validated by on-chain consensus. Bittensor’s consensus mechanism introduces latency (block finality: ~12 seconds) incompatible with synchronous gradient-based training, and is primarily targeted at inference marketplaces rather than pre-training runs. PRIME uses TOPLOC for compute verification without blockchain consensus, achieving verification latency under 100ms vs Bittensor’s 12-second block time.
Mathematical Framework: DiLoCo Outer Optimiser
- Pseudo-gradient formulation: Let θ^(t)_global be the globally synchronised model parameters at outer step t. Each worker i runs H inner optimiser steps, yielding local parameters θ^(t+1)_i. The pseudo-gradient for worker i is: Δ_i = θ^(t)_global − θ^(t+1)_i. This is the negative of the change in parameters induced by H local gradient steps—equivalent to an estimate of the gradient of the loss surface with respect to θ at a longer timescale than individual steps.
- Outer gradient aggregation: The outer pseudo-gradient is the all-reduce average: Δ_outer = (1/N) Σ_i Δ_i where N is the number of worker nodes. For uniform data distribution across workers, this is an unbiased estimator of the true full-batch pseudo-gradient.
- Nesterov outer update: θ^(t+1)_global = θ^(t)_global + α_outer × Δ_outer + β × m^(t) where m^(t) is the outer momentum term: m^(t+1) = β × m^(t) + α_outer × Δ_outer. Nesterov momentum (look-ahead momentum) was found empirically by Douillard et al. to outperform vanilla SGD or Adam for the outer optimiser, likely because it provides a smoothed estimate of the long-range gradient direction across multiple DiLoCo outer steps.
- Why the outer optimiser is not Adam: Adam’s adaptive per-parameter learning rates are well-suited to inner optimisation where gradients are noisy and heterogeneous across parameters. For the outer pseudo-gradient, which aggregates the net effect of 100 inner steps, the noise structure is different—pseudo-gradients are smoother and more consistent across outer steps, making Nesterov momentum more appropriate.
- int8 quantisation mathematics: Each pseudo-gradient tensor Δ ∈ R^d is quantised as: s = max(|Δ|) / 127; int8_Δ_i = round(Δ_i / s) ∈ {-127, …, 127}; transmitted as (s, int8_Δ). De-quantised: Δ̂_i = s × int8_Δ_i. Quantisation error: |Δ_i − Δ̂_i| ≤ s/2 = max(|Δ|)/254. For well-conditioned pseudo-gradients with typical max values ~0.01, quantisation error is < 5×10^−5—below numerical precision of fp16 training, making int8 quantisation effectively lossless.
- Convergence guarantee intuition: DiLoCo’s convergence theory (sketch from Douillard et al.): if the loss landscape is L-smooth and the inner optimiser runs H steps with learning rate η, the outer pseudo-gradient approximates the true gradient of a smoothed loss function averaged over H steps. Under standard SGD convergence analysis, the outer update makes progress proportional to ‖∇L‖² / (L × H × η)—meaning convergence rate degrades gracefully with H rather than catastrophically, as long as H is not too large relative to the loss curvature 1/L.
Data Pipeline and Training Data Quality
- FineWeb-Edu (55% of INTELLECT-1 data): HuggingFace’s FineWeb-Edu is a filtered subset of the Common Crawl web corpus, extracted by training an educational quality classifier on GPT-4 labels and filtering to documents scoring ≥3/5 on educational content. For INTELLECT-1, FineWeb-Edu provides the primary language modelling signal—conversational, instructional, and expository text from educational websites, textbooks, and academic sources.
- Stack v2 (20% code): The Stack v2 is a deduplication-focused software code dataset assembled by BigCode, containing source code from GitHub repositories across 600+ programming languages with near-duplicate removal and PII redaction. Code is a key modality for reasoning capability—training on code improves mathematical and structured reasoning in LLMs beyond code-specific tasks.
- DLCM (20% diverse language): Diverse Language Corpus Mixture combines multilingual web text, books, and news corpora to improve cross-lingual robustness and prevent English-only overfitting. Important for INTELLECT-1’s global compute contributor base, whose target use cases include non-English language assistance.
- OpenWebMath (5% mathematical text): OpenWebMath is a filtered corpus of mathematical text from the web, selected for high-quality mathematical exposition. The 5% weighting is intentionally small but provides concentrated signal for numerical and symbolic reasoning—disproportionate impact on mathematical benchmark performance relative to its data fraction.
- Data mixture rationale: The 55/20/20/5 split was determined empirically by ablation on smaller proxy models (1B parameters, 10B tokens), optimising for downstream benchmark performance on a held-out suite including Hellaswag, ARC, MMLU, HumanEval, and MATH. DiLoCo’s local optimiser means each node processes data from its local shard without cross-node data coordination; the global mixture target is achieved by pre-assigning data proportions to each node’s local dataset.
Organisational Philosophy and Governance
- Decentralised AI thesis: Prime Intellect’s founders articulate that concentration of AI training compute in a small number of entities (US hyperscalers, Chinese national labs, or a single frontier AI company) creates risks analogous to nuclear weapon monopoly—an entity controlling superintelligence training infrastructure could leverage that control for disproportionate geopolitical or economic power. Decentralised training infrastructure acts as a hedge against this concentration.
- Not anti-AI: Prime Intellect’s decentralisation argument is not that AI development should be slowed or prevented, but that the underlying infrastructure should be distributed—parallel to the argument that the internet’s resilience derives from its distributed routing, not a central switch. More compute contributors, more geographic diversity, and more open models reduces the ability of any single entity to control frontier AI.
- VitaDAO DeSci connection: Vincent Weisser’s co-founding of VitaDAO (the leading decentralised science DAO for longevity research, which has funded >$4M in research through tokenised IP-NFT mechanisms) informed Prime Intellect’s vision of tokenised AI ownership. VitaDAO demonstrates that community-funded, token-governed scientific organisations can successfully commission frontier research; Prime Intellect applies this model to AI training infrastructure.
- Open-source licensing strategy: Apache 2.0 licence (not GPL/copyleft) was chosen deliberately to maximise adoption by commercial entities who might build products on top of INTELLECT models and PRIME tooling. Copyleft would restrict commercial use and reduce the number of entities integrating Prime Intellect’s technology—counterproductive to the goal of building a broad decentralised AI ecosystem.
- Compute contributor incentives: Beyond payment for GPU time, Prime Intellect offers compute contributors: public acknowledgement in model papers and blogs (reputational value for hardware providers); early access to trained model weights; potential future ownership stakes under the planned tokenised protocol; and participation in a mission-driven project aligned with open AI development values.
- Differential technological development: The METAGENE-1 biosurveillance model exemplifies Prime Intellect’s “differential technological development” principle—prioritising AI applications that provide asymmetric defensive advantages (detecting novel pathogens earlier) without providing equivalent offensive capabilities (designing pathogens). The 512-token context and classification-only training ensure METAGENE-1 cannot be repurposed as a bioweapons design tool. This principle guides Prime Intellect’s selection of open science model candidates: applications where the public health benefit is high and misuse potential is constrained by architectural choice.
- Relationship to the AI safety debate: Prime Intellect’s decentralisation philosophy engages with but diverges from mainstream AI safety concerns about alignment and capabilities. The safety concern Prime Intellect addresses is structural—who controls AI training infrastructure—rather than behavioural (will the AI itself behave safely). Both concerns are legitimate; Prime Intellect’s position is that structural centralisation is an underappreciated risk that open-source and decentralised development directly mitigates, whilst acknowledging that behavioural alignment remains a separate challenge.
INTELLECT Series: Chronological Summary
- Timeline overview:
- July 2024: OpenDiLoCo open-source release; cross-continental DiLoCo training proof-of-concept at 1B scale.
- October 2024: INTELLECT-1 training run announcement; first globally distributed 10B training begins.
- November 2024: INTELLECT-1 base + instruct models released; 30 sponsors, 5 countries, 83% compute utilisation.
- December 2024: INTELLECT-1 technical report published (arXiv:2412.01152).
- January 2025: METAGENE-1 7B metagenomic foundation model released.
- February 2025: Series A $15M announced (Founders Fund lead, Menlo Ventures, angel investors).
- April 2025: INTELLECT-2 32B distributed RL launch; TOPLOC and SHARDCAST introduced.
- May 2025: INTELLECT-2 technical report (arXiv:2505.07291).
- November 2025: INTELLECT-3 106B MoE released; PRIME-RL framework open-sourced; arXiv:2512.16144.
- Parameter scale progression: 1B (OpenDiLoCo proof-of-concept) → 10B (INTELLECT-1) → 32B (INTELLECT-2) → 106B (INTELLECT-3): consistent ~3× scale increase per major run, tracking the broader industry trajectory of annual 10× parameter growth.
- Training paradigm progression: Supervised pre-training (OpenDiLoCo, INTELLECT-1) → distributed post-training RL (INTELLECT-2) → centralised large-scale RL for MoE (INTELLECT-3): systematic expansion across the full training pipeline.
- Geographic distribution progression: INTELLECT-1 spanned 5 countries on 3 continents (pre-training); INTELLECT-2’s rollout workers were globally distributed with centralised training; INTELLECT-3 was fully centralised. This reflects the evolving technical challenge: DiLoCo enables decentralised pre-training well, but decentralised RL at 100B+ scale requires additional engineering work targeted for INTELLECT-4.
- Open-source commitment: Every INTELLECT run has released model weights, training code, and training data under permissive open-source licences—a 100% open-source track record across the entire INTELLECT series as of mid-2026.
Technical Comparisons: PRIME vs Standard Training Infrastructure
- Memory efficiency: A standard data-parallel 10B training setup (no ZeRO) requires each GPU to hold full model + gradients + optimiser states = ~320 GB total per GPU—unfit on 80 GB H100. ZeRO-3 (FSDP2) reduces per-GPU memory to ~40 GB for 8-GPU node, enabling standard 8×H100 nodes without expensive 80GB multi-GPU instances. PRIME achieves the same memory efficiency as DeepSpeed ZeRO-3 within nodes whilst adding inter-node DiLoCo communication efficiency.
- Throughput numbers: INTELLECT-1 achieved approximately 2,800 tokens/second per 8×H100 node at 10B scale (batch size 2048, sequence length 2048). Across 10–14 active nodes (~80–112 H100s simultaneously), total training throughput reached approximately 28,000–39,000 tokens/second. To train on 1T tokens at this throughput requires approximately 7–10 hours per outer synchronisation cycle × 100 cycles ≈ 700–1,000 total training hours (confirmed by the ~40-minute inner step cycle from the technical report).
- Bandwidth requirements: With 4 Gbps inter-data-centre bandwidth (typical US inter-DC), an outer synchronisation transmitting 40 GB fp32 (10B model) would require 80 seconds = not 1 minute. PRIME’s int8 quantisation reduces outer sync data to 10 GB, transmissible in 20 seconds at 4 Gbps—confirmed by the <1 minute synchronisation time reported. US intra-region achieves 4 Gbps; transcontinental routes (US–Europe) achieve 1–2 Gbps, explaining the 83% vs 96% utilisation gap between global and US-only configurations.
- Scaling analysis: DiLoCo scales O(1) in communication bandwidth with number of nodes (same outer sync data size regardless of N nodes, because each node computes a fixed-size pseudo-gradient). Standard data-parallel all-reduce scales O(log N) in time with InfiniBand ring-allreduce, but O(1) is better when the constant in O(1) is low—which is the case for DiLoCo with int8 outer sync.
Key Risks and Limitations
- Convergence gap at larger H:
- DiLoCo’s theoretical analysis: increasing H (inner steps between outer syncs) degrades convergence rate.
- At H=100 (INTELLECT-1 configuration): degradation is minimal (<0.5 perplexity points vs synchronous baseline).
- At H=500 (original DeepMind DiLoCo paper): larger degradation but still acceptable for pre-training.
- At H=1,000+ (needed for very high-latency nodes like Starlink, latency >1 second per packet): convergence degradation may become problematic.
- Open research question for INTELLECT-4: whether improved outer optimisers (momentum adaptation, adaptive H) can maintain convergence at H=1,000+.
- Gradient staleness in distributed RL:
- INTELLECT-2 tolerates policy lag of up to 4 steps without measurable performance degradation.
- For larger models (32B+) or more complex reasoning tasks, the acceptable staleness threshold may decrease.
- Empirically determined constraint per INTELLECT-2 technical report; no theoretical guarantee that 4-step tolerance generalises to all task/model combinations.
- May limit the feasible degree of distribution for future RL runs if staleness tolerance shrinks below 1 step.
- Verifier gaming in TOPLOC:
- TOPLOC detects compute fraud via activation hashing—catches model weight modification, prompt substitution, precision changes.
- Attack surface: sophisticated adversary runs the correct model but manipulates final output text after computing a valid TOPLOC hash.
- Mitigation: fraudulent rollouts that pass TOPLOC but have incorrect content receive low rewards from the verifier (maths checker, code executor), contributing negligible gradient corruption.
- Residual risk: reward-agnostic attacks that produce plausible-looking but wrong outputs may pass both TOPLOC and low-quality reward models.
- Regulatory uncertainty:
- INTELLECT-1 had simultaneous nodes in the US, Europe, and Asia—three different regulatory jurisdictions for compute-intensive AI training.
- Emerging regulations: US export controls on A100/H100 GPUs already limit which countries can legally operate them; future compute licensing requirements could prohibit cross-border gradient exchange.
- Data localisation mandates (EU GDPR, China PIPL) may restrict which training datasets can be processed in which jurisdictions.
- Multi-jurisdictional compliance across 30+ independent compute contributors introduces significant legal coordination overhead.
- Economic sustainability:
- Prime Intellect as of mid-2026 remains VC-funded ($70.4M total) rather than operationally profitable.
- Revenue sources: GPU marketplace margin (typically 10–25% on compute reselling); potential future tokenised protocol fees; potential licensing of PRIME framework to enterprise users.
- Cost drivers: engineering team salaries (distributed systems and ML research engineers command >500K–$1M in GPU compute); infrastructure for marketplace orchestration.
- Long-term sustainability paths: marketplace volume reaching break-even scale (estimated >$10M/month GMV for margin-funded R&D); protocol token value appreciation; enterprise licensing of PRIME training framework.
Community and Ecosystem
- Compute contributor community: INTELLECT-1’s 30 sponsors demonstrate that a mix of commercial organisations (Hugging Face, SemiAnalysis), crypto/DeSci ecosystem players (Olas, Akash, Schelling AI), and individual GPU owners can collaborate in a single training run under a shared governance model (Prime Intellect as coordinator).
- Research community engagement: Prime Intellect publishes all INTELLECT training runs as peer-reviewed arXiv papers, participates in academic conferences (NeurIPS 2024, ICLR 2025), and maintains open-source repositories actively used by distributed ML researchers. The OpenDiLoCo repository attracted significant community contribution before deprecation.
- Hugging Face collaboration: Hugging Face is both a compute contributor (contributed H100 nodes to INTELLECT-1) and a distribution partner (INTELLECT model weights hosted on Hugging Face Hub). The Hugging Face Hub provides the open model weight distribution infrastructure that makes Prime Intellect’s open-source commitments practically accessible—without Hub, releasing 10B+ model weights would require self-hosted storage at significant cost.
- Akash Network integration: Akash Network’s blockchain-based permissionless GPU marketplace provides a constituency of small GPU providers (mining farms pivoting from cryptocurrency to AI compute) who cannot access centralised cloud contracts but can participate in Prime Intellect’s marketplace through Akash’s smart-contract-mediated bidding system. This integration bridges the DeSci/crypto compute ecosystem with mainstream ML infrastructure.
- Academic partnerships: METAGENE-1 demonstrates Prime Intellect’s model for academic collaboration—USC provides domain expertise (metagenomic sequencing, biological interpretation) and the Nucleic Acid Observatory provides wastewater sequencing data; Prime Intellect provides compute infrastructure, training expertise, and open-source distribution. This model could be replicated for climate, materials science, and drug discovery applications.
- Environments Hub community: PRIME-RL’s Environments Hub is designed as a community-extensible platform where academic researchers can contribute new RL task environments (maths verifiers, coding environments, scientific reasoning tasks, agentic tool-use benchmarks) and benefit from PRIME-RL’s distributed training infrastructure. This creates a flywheel: better environments attract more capable models; more capable models motivate more environment contributions.
INTELLECT-2: Detailed Technical Architecture
- Rollout worker specification: Any GPU with sufficient VRAM to run QwQ-32B inference (minimum 4×RTX 3090 at 24 GB each = 96 GB; alternative: 2×H100 at 80 GB each). Workers run vLLM for efficient autoregressive sampling. Each worker processes a batch of prompts in parallel (batch size 8–32 depending on available VRAM), generating 8 completions per prompt for GRPO group normalisation.
- Rollout data format: Each rollout consists of: prompt (maths problem or coding task, sourced from difficulty-filtered dataset with ≤75% solve rate); G=8 completions generated with temperature=0.9, top-p=0.95; TOPLOC hash signature; worker GPU type and precision; generation timestamp and policy version identifier.
- Centralised training node specification: GRPO training workers require much more compute than rollout workers—full FSDP2 training of QwQ-32B requires at minimum 8×H100 80GB (640 GB total for ZeRO-3 with full fp32 optimiser states). Training nodes receive verified rollout batches from TOPLOC validators, compute GRPO advantages, and perform gradient steps.
- GRPO advantage computation: For each rollout group (G=8 completions on a single prompt): r_i = reward from verifier for completion i; A_i = (r_i − mean_{j=1}^G r_j) / (std_{j=1}^G r_j + ε). Advantage normalisation within group removes reward scale sensitivity—a reward of 1 for a correct answer has the same gradient contribution regardless of whether other completions also scored 1 or 0.
- KL regularisation: GRPO loss includes a KL divergence penalty between current policy π and frozen reference policy π_ref: L_GRPO = -E[A_i × log π(completion_i | prompt_i)] + β_KL × KL(π || π_ref). KL coefficient β_KL=0.01 prevents the policy from collapsing to reward hacking—the model cannot drastically change its output distribution from the initial QwQ-32B checkpoint in a way that exploits verifier quirks without paying a KL penalty.
- Difficulty filtering rationale: Including easy problems (>75% solve rate) in RL training creates batches where all G completions have reward=1 (all correct), yielding zero advantage variance. Zero advantage variance means zero gradient—these batches contribute no learning signal. Filtering to ≤75% solve rate ensures at least some completions are incorrect, maintaining non-zero advantage variance and efficient gradient utilisation.
- Online advantage filtering: During training, batches where advantage variance falls below threshold σ_min² are discarded. This filters: (a) remaining trivially easy problems where all completions happen to be correct; (b) reward hacking patterns where all completions find a spurious solution that the verifier accepts but the reasoning is invalid; (c) reward collapsing where all completions fail equally (all reward=0, all advantage=0).
Hardware Ecosystem and GPU Specifications
- NVIDIA H100 (80GB SXM5): Primary training GPU for all INTELLECT runs. 80 GB HBM2e memory; 3,958 TFLOPS fp16 tensor core throughput; 900 GB/s HBM2e bandwidth; 600 GB/s NVLink-4 bandwidth per GPU (full NVSwitch fabric: 900 GB/s bidirectional). Used in DGX H100 8-GPU nodes with NVSwitch full connectivity.
- NVIDIA H200 (141GB SXM5): Primary training GPU for INTELLECT-3. 141 GB HBM3e memory (75% more than H100); 4,915 TFLOPS fp8 throughput; 3.35 TB/s HBM3e bandwidth (60% more than H100). Larger memory enables larger batch sizes and less frequent gradient checkpointing, improving effective throughput for 106B MoE training.
- RTX 3090 (24GB): Minimum contributor GPU for INTELLECT-2 rollout workers. 24 GB GDDR6X; 142 TFLOPS fp16; sufficient for QwQ-32B inference in 4-way tensor parallelism (VRAM usage ~96 GB across 4 GPUs). 4×RTX 3090 cost approximately $4,000–6,000 USD in 2025, making INTELLECT-2 participation accessible to individual researchers.
- Heterogeneous hardware coordination: PRIME’s ElasticDeviceMesh handles heterogeneous inner-node hardware (some contributors may use 8×H100, others 8×A100 80GB, others 4×A100 40GB). The inner FSDP2 sharding adapts to available VRAM; the outer DiLoCo sync is hardware-agnostic (pseudo-gradients have identical structure regardless of which GPU computed them). This hardware heterogeneity tolerance is essential for aggregating diverse supply from the marketplace.
INTELLECT Series Quick-Reference Specifications
| Run | Parameters | Training Type | Nodes/GPUs | Countries | Compute Utilisation | Release |
|---|---|---|---|---|---|---|
| OpenDiLoCo | ~1B | Pre-training (DiLoCo) | 2 continents, 3 countries | 3 | 90–95% | Jul 2024 |
| INTELLECT-1 | 10B | Pre-training (DiLoCo + FSDP2) | 30 sponsors, 80–112 H100s | 5 | 83% global / 96% US | Nov 2024 |
| INTELLECT-2 | 32B (QwQ) | Distributed RL (async GRPO) | Consumer GPUs (min 4×RTX 3090) | Global | N/A (async) | Apr 2025 |
| INTELLECT-3 | 106B MoE | Centralised RL | 512×H200 (64 nodes) | 1 (centralised) | N/A (centralised) | Nov 2025 |
| METAGENE-1 | 7B | Pre-training (genomics) | Prime Intellect + USC | N/A | N/A | Jan 2025 |
Communication Efficiency Comparison Table
| Training Method | Communication per Step | Reduction vs Baseline |
|---|---|---|
| Standard FP32 data-parallel | 40 GB (10B model) | 1× (baseline) |
| FP16 data-parallel | 20 GB | 2× |
| DiLoCo FP32 (H=100) | 400 MB effective avg | 100× |
| DiLoCo INT8 (H=100) | 100 MB effective avg | 400× |
| DiLoCo INT8 (H=500) | 20 MB effective avg | 2,000× |
- Interpretation: At 4 Gbps internet bandwidth per node, standard FP32 data-parallel all-reduce requires 80 seconds per step—making step-by-step training completely infeasible over the internet. DiLoCo INT8 at H=100 requires 0.2 seconds per effective step (100 MB / 500 MB/s)—feasible even on modest internet connections. At H=500, PRIME requires only 40 milliseconds per effective step, within Ethernet switch latency territory.
Funding and Investment Summary
| Round | Date | Amount | Lead Investors | Strategic Notes |
|---|---|---|---|---|
| Seed | April 2024 | $5.5M | CoinFund, Distributed Global | Former OpenAI founding member angels; crypto/DeSci alignment |
| Series A | February 2025 | $15M | Founders Fund (Peter Thiel), Menlo Ventures | Tech/libertarian alignment with decentralisation thesis; angels include Karpathy, Delangue, Tri Dao |
| Subsequent | 2025–2026 | $49.9M (est.) | Undisclosed | Total to $70.4M per Sacra/Crunchbase data |
- Investor alignment: CoinFund (crypto-native VC) and Distributed Global (decentralised infrastructure focus) in the seed round signal Prime Intellect’s positioning within the blockchain/DeSci ecosystem. Founders Fund’s Series A investment—Peter Thiel’s fund known for contrarian bets on decentralised and disruptive technology—validates the thesis at a larger capital scale. The investor mix across seed and Series A reflects Prime Intellect’s dual positioning as both a crypto-ecosystem project and a mainstream AI infrastructure company.
Metadata
| Field | Value |
|---|---|
| Domain correction | infrastructure → artificial-intelligence (Prime Intellect is an AI research/training organisation; original domain was misclassified) |
| IRI corrected | http://narrativegoldmine.com/artificial-intelligence#PrimeIntellect |
| URI corrected | urn:visionclaw:concept:artificial-intelligence:prime-intellect |
| same-as corrected | urn:visionclaw:concept:artificial-intelligence:prime-intellect |
| owl-class corrected | artificial-intelligence:PrimeIntellect |
| legacy-term-id | AI-2047 |
| Research cache | _enrich/research-cache/Prime Intellect.json |
| Worker model | claude-sonnet-4-6 |
| Phase | 6 |
Provenance
- arXiv:2412.01152 — INTELLECT-1 Technical Report (Jaghouar et al., December 2024)
- arXiv:2407.07852 — OpenDiLoCo (Jaghouar et al., August 2024)
- arXiv:2505.07291 — INTELLECT-2 Technical Report (Prime Intellect Team, May 2025)
- arXiv:2512.16144 — INTELLECT-3 Technical Report (Prime Intellect Team, November 2025)
- arXiv:2501.02045 — METAGENE-1 (Senghaas et al., January 2025)
- arXiv:2604.21428 — Decoupled DiLoCo (Google DeepMind, April 2025)
- arXiv:2501.12948 — DeepSeek-R1 / GRPO (DeepSeek-AI, January 2025)
- https://www.primeintellect.ai/blog/intellect-1 — INTELLECT-1 launch blog
- https://www.primeintellect.ai/blog/intellect-1-release — INTELLECT-1 release
- https://www.primeintellect.ai/blog/intellect-2 — INTELLECT-2 announcement
- https://www.primeintellect.ai/blog/intellect-2-release — INTELLECT-2 release
- https://www.primeintellect.ai/blog/intellect-3 — INTELLECT-3 release
- https://www.primeintellect.ai/blog/opendiloco — OpenDiLoCo blog
- https://www.primeintellect.ai/blog/metagene — METAGENE-1 blog
- https://www.primeintellect.ai/blog/compute — Prime Intellect Compute marketplace
- https://github.com/PrimeIntellect-ai/OpenDiloco — OpenDiLoCo GitHub
- https://github.com/PrimeIntellect-ai/prime-rl — PRIME-RL GitHub
- https://deepmind.google/research/publications/57039/ — DiLoCo original (Douillard et al.)
- https://deepmind.google/blog/decoupled-diloco/ — Decoupled DiLoCo (Google DeepMind)
- https://www.prnewswire.com/news-releases/prime-intellect-secures-5-5m-in-seed-funding — Seed round press release
- https://fortune.com/2024/04/23/coinfund-distributed-global-prime-intellect — Fortune seed coverage
- https://akash.network/blog/prime-intellect-integrates-permissionless-akash-gpus/ — Akash integration
- https://nebius.com/customer-stories/prime-intellect — Nebius GB200 partnership
- https://sacra.com/c/prime-intellect/ — Sacra funding analysis
- https://huggingface.co/PrimeIntellect/INTELLECT-3 — INTELLECT-3 model card
- https://metagene.ai/ — METAGENE-1 site
- https://www.cognitiverevolution.ai/distributed-training-decentralized-ai-prime-intellects-master-plan — Podcast interview
- https://medium.com/theta-network/imperial-college-london-to-use-theta-edgecloud-in-its-ai-research — Imperial/Theta EdgeCloud
- domain-correction: infrastructure → artificial-intelligence