The Coe Lab
← Back to Blog

GigaToken: The 1000x Faster Tokenizer Quietly Rewriting AI's Data Pipeline

By July 23, 20266 min read
AItokenizationopen-sourceinfrastructuremachine-learning
A technology notebook with visual symbols for AI, cybersecurity, infrastructure, and automation

A new open-source tokenizer called GigaToken is hitting GB/s throughput — up to 1000x faster than HuggingFace's tokenizers. It could fundamentally change how AI labs preprocess the trillions of tokens that train today's frontier models.

Tokenization is the unglamorous plumbing of the AI world. Every language model — from GPT-5.6 to DeepSeek V4 — depends on a tokenizer to chop raw text into the discrete units the model actually processes. It happens before training, before inference, before any of the magic. And until now, almost nobody talked about it. Tokenizers are the invisible foundation of every language model, and like most foundations, they get taken for granted until something goes wrong.

That changed this week when a project called GigaToken landed on Hacker News and rocketed to over 500 points. Its claim? Tokenize text at gigabytes per second — up to 1000x faster than HuggingFace's widely-used tokenizers library, and hundreds of times faster than OpenAI's tiktoken. Both of those are already written in multithreaded Rust. GigaToken is too, but it's doing something fundamentally different with the architecture of the tokenization pipeline.

What Makes GigaToken Different

Most tokenizers today follow a familiar pattern: read text into memory, run byte-pair encoding (BPE) or SentencePiece, emit token IDs. HuggingFace's tokenizers and tiktoken both do this, and they're fast — tens of megabytes per second on a single core. For most use cases, this is perfectly adequate. You tokenize your dataset once, cache the results, and move on to the interesting parts of model development.

GigaToken rethinks the entire pipeline from the ground up. The key insight is that tokenization is embarrassingly parallel — the kind of problem where you can divide the work across many independent workers with minimal coordination. When you're processing a 12 GB text corpus, you don't need to load it all into memory and process it sequentially. GigaToken reads data directly from disk in chunks, tokenizes each chunk in parallel across all available cores, and streams the results out — all in Rust, with minimal Python overhead. The result is a pipeline that saturates your disk I/O and CPU bandwidth rather than bottlenecking on a single thread.

The result is throughput numbers that sound like a typo:

  • 24.5 GB/s for GPT-2 tokenization on a 144-core AMD EPYC system — fast enough to tokenize the entire English Wikipedia in under a second
  • 8.8 GB/s on an Apple M4 Max (16 cores) — faster than most NVMe SSDs can sustain reads
  • 6.3 GB/s on an AMD Ryzen 7 9800X3D (16 threads) — bringing near-instantaneous tokenization to desktop hardware
  • Up to 1,268x faster than HuggingFace tokenizers on the M4 Max — not an incremental improvement but an order-of-magnitude leap
  • To put that in perspective: tokenizing the entire OpenWebText corpus (11.9 GB) takes roughly half a second on the EPYC system with GigaToken, versus several minutes with HuggingFace's tokenizer. For a research lab iterating on training data, the difference between half a second and five minutes per tokenization pass fundamentally changes the workflow.

    Why This Actually Matters

    It's tempting to dismiss tokenization speed as an engineering footnote. Frontier model training runs cost tens of millions of dollars and take weeks. What does a few minutes of tokenization save? Quite a lot, actually, and the reason has to do with how modern AI development actually works in practice.

    Modern AI labs don't tokenize once and move on. They tokenize constantly — for data exploration, for filtering, for evaluation pipelines, for synthetic data generation, for re-training on new data mixes, for data deduplication, for quality scoring, for language identification, for toxic content filtering. Every iteration of the data pipeline involves tokenization. When you're iterating on a training corpus, shaving tokenization from hours to seconds changes how you work. Engineers can try a new data mix, tokenize it, and start a training run in minutes instead of overnight. The tight feedback loop enables more experimentation, which leads to better data curation, which leads to better models.

    There's also a memory story that matters for large-scale operations. GigaToken's streaming approach means you can tokenize datasets that don't fit in RAM — a real problem when your corpus is measured in terabytes, not gigabytes. HuggingFace's tokenizers need to hold the text in memory, which means you need a machine with enough RAM to hold your entire dataset, or you need to chunk it yourself. GigaToken just streams through it, reading from disk and writing token IDs in a continuous flow. This enables tokenization of arbitrarily large datasets on modest hardware.

    Broad Model Support, Drop-In Replacement

    One of GigaToken's strongest selling points is its compatibility layer. It works with virtually every major tokenizer in use today, making adoption essentially frictionless for existing codebases:

  • GPT-2 / GPT-OSS (OpenAI's open models) — the most widely used BPE tokenizer in the research community
  • Llama 3, 3.1, 3.2, 3.3, and Llama 4 — Meta's family of tokenizers used in the most popular open-weights models
  • Qwen 2, 2.5, 3, 3.5, and 3.6 — Alibaba's tokenizer family, widely used in multilingual and coding models
  • DeepSeek V3, R1, and V4 — the tokenizers behind some of the most efficient reasoning models
  • GLM 4 and GLM 5 — Zhipu AI's tokenizers used in the GLM model family
  • Phi-4, Mistral, Gemma, Kimi K2, and more — covering the rest of the major model families
  • You can use it in compatibility mode with HuggingFace Tokenizers or tiktoken, which means existing code barely needs to change — swap your import and you're done. The tradeoff is that compatibility mode adds some overhead — you get maybe 10-100x speedup instead of the full 1000x. For maximum throughput, GigaToken's native API reads files directly from Rust, bypassing Python entirely. This gives you the full performance benefit but requires slightly more code changes. The tradeoff between ease of adoption and maximum performance is well-calibrated for real-world use.

    The Catch: SentencePiece Tokenizers Lag Behind

    GigaToken isn't uniformly fast across all tokenizers. The results show a clear divide that matters for some users. BPE-based tokenizers (GPT-2, Llama, Qwen, DeepSeek) hit the spectacular numbers — the 1000x speedups that grab headlines. But SentencePiece-based tokenizers (Gemma, Mistral, CodeLlama, TinyLlama) see much more modest gains — 7x to 22x. That's still a meaningful improvement, but it's not the headline-grabbing 1000x.

    The developer notes that SentencePiece tokenizers are "not well optimized" in GigaToken yet, which suggests there's room for improvement. The SentencePiece algorithm has different internal mechanics — it uses a unigram language model rather than BPE's merge-based approach, and the data structures required for efficient lookup are different. But the BPE results alone cover the vast majority of today's frontier models, so for most users this caveat is academic. For teams working with Gemma or Mistral, the 7-22x speedup is still worth the migration effort.

    Implications for the AI Ecosystem

    GigaToken arrives at a moment when AI labs are drowning in data. Training runs increasingly involve trillions of tokens, and data pipelines are becoming a competitive bottleneck. The quality of your training data matters more than the size of your model — this has been demonstrated repeatedly in 2026 — and data quality is a function of how many iterations you can run. If you can tokenize 10x faster, you can iterate on data curation 10x more often — and data quality is arguably the biggest lever in model performance today.

    It also levels the playing field. Smaller labs and independent researchers who can't afford massive preprocessing clusters can now tokenize at speeds that were previously the domain of well-funded AI companies with dedicated infrastructure teams. A researcher with a single workstation can process datasets at speeds that would have required a cluster just months ago. This democratization of data processing capability could enable breakthroughs from unexpected places — university labs, independent researchers, small startups — that have been bottlenecked on data infrastructure rather than ideas.

    And there's something satisfying about progress happening at the most fundamental layer of the AI stack. We've gotten used to breakthroughs in model architecture, training algorithms, and reasoning capabilities. But sometimes the biggest wins come from revisiting the boring infrastructure that everything else is built on — and making it 1000x faster. GigaToken is a reminder that the AI stack has many layers, and innovation at any layer can unlock new possibilities at every layer above it.

    GigaToken is open source, available on GitHub, and installable via pip. If you're working with language models at any scale — from fine-tuning experiments to trillion-token pretraining runs — it's worth a serious look. The installation takes seconds, the API is familiar, and the performance gains are immediate and measurable. In a field where we chase 2% improvements in benchmark scores, a 1000x improvement in pipeline throughput is a rare and welcome gift.

    Related Posts

    Claude Haiku 5.5: Why Cheap AI Changes Agent Architecture

    Claude Haiku 5.5 cuts small-model costs dramatically while adding serious agent skills. Here is why routing, caching, and architecture now matter more than model size.

    Oct 8, 2026• 10 min

    When the Registry Fell: How Hijacked Country Domains Became the New Attack Vector for Counterfeit TLS Certificates

    Attackers compromised three country-code top-level domain registries to mint fraudulent HTTPS certificates for Google and other major services. The incident exposes a structural weakness in the web's trust infrastructure that no browser alone can fix.

    Oct 7, 2026• 8 min read

    Cloudflare's Web Search API: When the Edge Network Became the Search Engine for AI Agents

    Cloudflare's new Web Search API gives AI agents real-time web search through AI Gateway with three providers, unified billing, and zero-config Workers integration. Here is what developers need to know.

    Oct 6, 2026• 8 min read