Menu

Post image 1
Post image 2
1 / 2
0

GitHub - marcelroed/gigatoken: Language model tokenization at GB/s

Hacker News·3 months ago
#LrKukuZn
Reading 0:00
15s threshold

~1000x faster than HuggingFace's tokenizers, drop-in replacement. Tokenize your text data at GB/s! Note that both HF tokenizers and tiktoken are already running multithreaded Rust! What is Gigatoken? Gigatoken is the fastest tokenizer for language modeling. It supports a wide range of CPU hardware, and nearly all commonly used tokenizers. See the Benchmarks section for detailed throughput numbers across tokenizers and CPUs. Installation Usage Gigatoken can be used with its own API, or in compatibility mode with HuggingFace Tokenizers or Tiktoken. Compatibility Mode (Easiest) import gigatoken as gt # Minimum change from existing HuggingFace tokenizers usage (compatibility mode) hf_tokenizer = ... tokenizer = gt . Tokenizer ( hf_tokenizer ). as_hf () # tokenizer can be used in the same contexts as hf_tokenizer tokens = tokenizer . encode_batch ([ "This is a test string" , "And here is another" ]) # OR with tiktoken tiktokenizer = ... tokenizer = gt . Tokenizer ( tiktokenizer ).…

Continue reading — create a free account

Join HashtagPLUS to read full articles, follow hashtags, vote, and join the conversation.

Read More