Tag
A tool that enables autocorrection for typing without using spacebars, available as an open-source project on GitHub.
TXR is an original, new programming language designed for convenient data munging, offering tools for text processing and data manipulation with resources like documentation, downloads, and community support.
An experiment tests a text-only AI model's ability to recognize shapes by converting images to Unicode braille text, achieving modest accuracy above chance but with reliable confidence estimates.
Milliseconds.ai is a fast AI API for text and image processing, featuring a small model called decision-machine-1, with competitive pricing and a free tier.
semfont is a library that automatically applies typographic styles to text based on word scores for valence, salience, surprise, and certainty to enhance readability.
Desert Ant Labs offers small, specialized AI models for speech, text, and vision that run offline on devices, with an SDK for easy integration and no per-use costs.
This article explains how to use the m4 macro processor to reduce repetitive code in OpenBSD's httpd configuration by creating custom macros.
Google DeepMind announces improvements to their AI model, enhancing its ability to understand complex phone numbers, postal codes, and order IDs in noisy environments, along with features like removing filler words and recognizing custom vocabulary.
Introduces grapheme-kit, an open-source Python library that extends lexical distance and evaluation metrics to operate on grapheme clusters instead of Unicode code points, with improved processing for Tamil and Sinhala.
Introduces 7 GitHub projects for removing AI writing traces, covering both Chinese and English texts, including tools like qu-ai-wei and Humanizer, to help users write articles that sound more human.
A 1997 tutorial explaining how to use sed and sort to create formatted indexes for books from raw term-page number lists.
Gigatoken is an open-source tokenizer that achieves up to 1000x speedup over HuggingFace tokenizers and 100x over Tiktoken, using SIMD and caching optimizations. It supports drop-in replacement for existing tokenizer APIs.
Mojibake is a self-contained Unicode 17 library for C11/C++17, providing normalization, case conversion, and character database functions with zero dependencies.
ColibotAI is an on-device AI tool that translates, summarizes, and explains any text without needing internet connection.
This repository compresses 201GB of text down to 6GB with no accuracy loss, making it 97% smaller than vector databases. It runs locally and offers a drop-in MCP for Claude, fully open source and private.