text-processing

Tag

Cards List
#text-processing

Jev-powered autocorrection

Lobsters Hottest ↗ · 6d ago Cached

A tool that enables autocorrection for typing without using spacebars, available as an open-source project on GitHub.

0 favorites 0 likes
#text-processing

TXR: An Original, New Programming Language for Convenient Data Munging

Hacker News Top ↗ · 2026-09-21 Cached

TXR is an original, new programming language designed for convenient data munging, offering tools for text processing and data manipulation with resources like documentation, downloads, and community support.

0 favorites 0 likes
#text-processing

@BenjaminDEKR: Can Jev understand basic shapes? Kind of. Jev is text-only. No image input at all. So I cheated. Render a shape to a 64…

X AI KOLs Following ↗ · 2026-09-20 Cached

An experiment tests a text-only AI model's ability to recognize shapes by converting images to Unicode braille text, achieving modest accuracy above chance but with reliable confidence estimates.

0 favorites 0 likes
#text-processing

Milliseconds.ai

Product Hunt ↗ · 2026-09-20 Cached

Milliseconds.ai is a fast AI API for text and image processing, featuring a small model called decision-machine-1, with competitive pricing and a free tier.

0 favorites 0 likes
#text-processing

A font that reads what you wrote

Hacker News Top ↗ · 2026-09-20 Cached

semfont is a library that automatically applies typographic styles to text based on word scores for valence, salience, surprise, and certainty to enhance readability.

0 favorites 0 likes
#text-processing

Desert Ant Labs

Product Hunt ↗ · 2026-09-09 Cached

Desert Ant Labs offers small, specialized AI models for speech, text, and vision that run offline on devices, with an SDK for easy integration and no per-use costs.

0 favorites 0 likes
#text-processing

Repeating Ourselves Less with M4

Lobsters Hottest ↗ · 2026-08-31 Cached

This article explains how to use the m4 macro processor to reduce repetitive code in OpenBSD's httpd configuration by creating custom macros.

0 favorites 0 likes
#text-processing

@GoogleDeepMind: Here’s what’s new: It’s better at understanding complex phone numbers, postal codes, and order IDs – even in noisy envi…

X AI KOLs ↗ · 2026-08-26 Cached

Google DeepMind announces improvements to their AI model, enhancing its ability to understand complex phone numbers, postal codes, and order IDs in noisy environments, along with features like removing filler words and recognizing custom vocabulary.

0 favorites 0 likes
#text-processing

grapheme-kit: Grapheme-Level Metrics and Text Processing for Multilingual NLP

arXiv cs.CL ↗ · 2026-07-27 Cached

Introduces grapheme-kit, an open-source Python library that extends lexical distance and evaluation metrics to operate on grapheme clusters instead of Unicode code points, with improved processing for Tamil and Sinhala.

0 favorites 0 likes
#text-processing

@0xCheshire: I researched how to remove the AI tone and found 7 different GitHub projects: for Chinese rewriting, English rewriting, technical writing, and complete writing workflows, each with corresponding options. 1. qu-ai-wei (299 stars) Chinese de-AI Skill, which handles clichés, abstract expressions...

X AI KOLs Timeline ↗ · 2026-07-26 Cached

Introduces 7 GitHub projects for removing AI writing traces, covering both Chinese and English texts, including tools like qu-ai-wei and Humanizer, to help users write articles that sound more human.

0 favorites 0 likes
#text-processing

Using sed to make indexes for books (1997)

Hacker News Top ↗ · 2026-07-23 Cached

A 1997 tutorial explaining how to use sed and sort to create formatted indexes for books from raw term-page number lists.

0 favorites 0 likes
#text-processing

Gigatoken: A new open source tokenizer ~100x faster than Tiktoken, -500-1000x faster than Huggingface

Reddit r/LocalLLaMA ↗ · 2026-07-21 Cached

Gigatoken is an open-source tokenizer that achieves up to 1000x speedup over HuggingFace tokenizers and 100x over Tiktoken, using SIMD and caching optimizations. It supports drop-in replacement for existing tokenizer APIs.

0 favorites 0 likes
#text-processing

Show HN: Mojibake – a low-level Unicode library written in C

Hacker News Top ↗ · 2026-07-16 Cached

Mojibake is a self-contained Unicode 17 library for C11/C++17, providing normalization, case conversion, and character database functions with zero dependencies.

0 favorites 0 likes
#text-processing

ColibotAI

Product Hunt ↗ · 2026-06-09

ColibotAI is an on-device AI tool that translates, summarizes, and explains any text without needing internet connection.

0 favorites 0 likes
#text-processing

@HowToAI_: This repo shrinks 201GB of text down to 6GB without losing any accuracy. → 97% smaller than vector DBs → Runs locally →…

X AI KOLs Timeline ↗ · 2026-05-15 Cached

This repository compresses 201GB of text down to 6GB with no accuracy loss, making it 97% smaller than vector databases. It runs locally and offers a drop-in MCP for Claude, fully open source and private.

0 favorites 0 likes
← Back to home

Submit Feedback