tokenizer

Tag

Cards List
#tokenizer

Claude Sonnet 5 Could Be Released Later Today, But May Not Be Better Than Opus 4.8

Reddit r/singularity · 2026-06-30

Claude Sonnet 5 may be released later today, featuring a new tokenizer, high-resolution vision, and marketed as Sonnet-priced near-Opus performance, but it may not exceed Opus 4.8.

0 favorites 0 likes
#tokenizer

BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation

arXiv cs.AI · 2026-06-20 Cached

Introduces BrainG3N, a dual-purpose tokenizer for 3D brain MRI latent diffusion using a frozen masked autoencoder encoder for clinically informative embeddings and a CNN decoder for reconstruction, achieving state-of-the-art performance on a 23-task benchmark and enabling controllable generation and longitudinal forecasting.

0 favorites 0 likes
#tokenizer

quicktok: a faster tokenizer (exact and byte-identical with tiktoken) [P]

Reddit r/MachineLearning · 2026-06-16

quicktok is a fast and exact BPE tokenizer in C++ that is byte-identical with tiktoken, achieving 2–11x speedup over existing alternatives. It supports cl100k, o200k, GPT-OSS, Llama-3, and Qwen2.5/3 encoders.

0 favorites 0 likes
#tokenizer

@PierceZhang34: Train a Small Model in 10 Seconds! First Look at the LLM Training Tool: http://llm.istanbul Recently discovered a super fun open-source style tool website — http://llm.istanbul, which claims to be a WebGPU LLM Workbench, meaning it fully...

X AI KOLs Timeline · 2026-06-12 Cached

Introduces llm.istanbul, a WebGPU LLM workbench that lets you train small models, train tokenizers, and generate text entirely in the browser, no server required, fully local.

0 favorites 0 likes
#tokenizer

Time Series as Language: A Universal Tokenizer for General-Purpose Time Series Foundation Models

arXiv cs.LG · 2026-06-10 Cached

Introduces UniTok, a universal tokenizer that transforms continuous time series into discrete tokens, and UniTok-FM, a foundation model pretrained via next-token prediction that enables zero-shot and prompt-boosted forecasting as well as few-shot generation and classification through training-free in-context inference.

0 favorites 0 likes
#tokenizer

MiniCPM5 1B - what is it?

Reddit r/LocalLLaMA · 2026-06-01

MiniCPM5-1B is a new small language model from OpenBMB, apparently built from scratch with its own tokenizer and distinct behavior, generating excitement as a capable 1B model.

0 favorites 0 likes
#tokenizer

Add MiniCPM5 tokenizer support by zhangtao2-1 · Pull Request #23384 · ggml-org/llama.cpp

Reddit r/LocalLLaMA · 2026-05-27 Cached

This pull request adds tokenizer support for MiniCPM5 to llama.cpp, extending the tool's compatibility with the MiniCPM family of models.

0 favorites 0 likes
#tokenizer

[NEW] Supra-50M Released!

Reddit r/LocalLLaMA · 2026-05-22

SupraLabs released Supra-50M, a compact 50M-parameter causal language model with base and instruct versions, trained on 20B tokens from fineweb-edu, achieving competitive benchmarks against larger models like GPT-2 and SmolLM.

0 favorites 0 likes
#tokenizer

ztok — a fast multithreaded tokenizer in Zig that loads tiktoken / HF / SentencePiece and is 2–5× faster

Reddit r/LocalLLaMA · 2026-05-22

ztok 是一个用 Zig 编写的高性能多线程分词器库,支持多种格式(tiktoken、HF、SentencePiece 等),速度比现有方案快 2–5 倍,适用于 RAG 分块和数据集分词。

0 favorites 0 likes
#tokenizer

@lvwerra: We are releasing Carbon: a crazy fast DNA model Carbon is 275x faster than the next best model. So fast you can process…

X AI KOLs Following · 2026-05-19 Cached

HuggingFace releases Carbon, a DNA model that is 275x faster than the previous state-of-the-art (Evo2), enabling processing of the entire human genome on a single GPU in under two days. The model uses a unique tokenizer that splits sequences into 6-base chunks while maintaining single-base resolution, and comes with an interactive demo.

0 favorites 0 likes
#tokenizer

Number-aware embeddings

Reddit r/LocalLLaMA · 2026-05-19

A technique to make embedding models aware of number ordering by overriding tokenizer and MLM fine-tuning, achieving 59% accuracy on number sorting benchmarks.

0 favorites 0 likes
#tokenizer

The biggest AI breakthrough in medicine & drug discovery

Reddit r/singularity · 2026-05-14 Cached

MAML is a novel multi-modal AI model that unifies understanding of chemistry, genetics, and proteins, outperforming specialized models on 11 drug discovery benchmarks, promising to accelerate pharmaceutical research and improve success rates.

0 favorites 0 likes
#tokenizer

Attacks On Data Centers, Qwen3.5 In All Sizes, DeepSeek’s Huawei Play, Apple’s Multimodal Tokenizer

The Batch · 2026-03-20 Cached

Andrew Ng's newsletter covers recent AI developments including attacks on data centers, the release of Qwen3.5 in various sizes, DeepSeek's collaboration with Huawei, and Apple's multimodal tokenizer, alongside reflections on AI-driven job uncertainty and geopolitical risks.

0 favorites 0 likes
#tokenizer

shiyu-coder/Kronos

GitHub Trending (daily) · 2026-05-14 Cached

Kronos is an open-source foundation model for financial K-line sequences, trained on data from over 45 global exchanges. It uses a specialized tokenizer and a decoder-only Transformer, and has been accepted at AAAI 2026.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback