training

Tag

Cards List
#training

@fuzzsociety_org: Linux Kernel Exploitation will be released soon, don't forget to reserve your spots at a discounted price

X AI KOLs Timeline ↗ · 2026-09-17 Cached

An announcement for the upcoming release of Linux Kernel Exploitation, with discounted reservation spots available.

0 favorites 0 likes
#training

Update : Small model + Engram

Reddit r/LocalLLaMA ↗ · 2026-09-17

The author provides an update on building a small 2B parameter AI model with an Engram component, trained on 15m tokens from Wikipedia to achieve surprising coherence, with plans for an Apache 2.0 open-source release.

0 favorites 0 likes
#training

@Kay2289123: That ghost story about AI development slowing down—read it, have a laugh, and move on; don't fool even yourself. Grok4.…

X AI KOLs Timeline ↗ · 2026-09-16 Cached

The article refutes the idea that AI development is slowing down, highlighting that Grok4.8, a 2.5T parameter model, has finished training and that major tech companies are still actively training AI models.

0 favorites 0 likes
#training

Landing the Space Shuttle – A Flying Machine and the Thrill of a Lifetime

Hacker News Top ↗ · 2026-09-16 Cached

An astronaut recounts the thrill of landing the Space Shuttle Atlantis, detailing the rigorous training and precision required for the dead-stick descent from Mach 1 to touchdown.

0 favorites 0 likes
#training

@HuggingPapers: ZGCM-1: 7B model, frontier-level math and search A fully open 7B LLM, now on Hugging Face. Trained from scratch with FP…

X AI KOLs Timeline ↗ · 2026-09-15 Cached

ZGCM-1 is a fully open 7B LLM trained from scratch with FP8, MDP mid-training, and AI4AI, achieving performance comparable to Qwen3-235B on math and agentic search tasks.

0 favorites 0 likes
#training

@kentcdodds: This guy is always talking about the next next thing before we have the next thing

X AI KOLs Timeline ↗ · 2026-09-14 Cached

Elon Musk announces that Grok 4.8, a 2.5 trillion parameter AI model trained with a new C++ software stack, will finish training and start reinforcement learning this week, while @kentcdodds comments on Musk's tendency to discuss future developments early.

0 favorites 0 likes
#training

Disentangling Representation Evolution in Transformers through Directional Decomposition

Hugging Face Daily Papers ↗ · 2026-09-14 Cached

This paper decomposes transformer representation updates into parallel and perpendicular components to study evolution geometry, linking it to editing robustness, compression diagnosis, and training improvements.

0 favorites 0 likes
#training

Pacing the Frontier – Tahuna: AI Training Infrastructure, Now Open Source [P]

Reddit r/MachineLearning ↗ · 2026-09-13

Tahuna is an open-source AI training infrastructure designed for small teams to train models, run inference, orchestrate GPUs, and experiment with autonomous research, featuring reproducible runs and tools like Hillclimb.

0 favorites 0 likes
#training

Is there still strong interest in a dense 9b model?

Reddit r/LocalLLaMA ↗ · 2026-09-13

A developer has created a dense 9.4B parameter AI model with technical enhancements like Engram tables and AttnRes modeling, aiming for open-source release and seeking community interest for further training and deployment.

0 favorites 0 likes
#training

Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training

Hugging Face Daily Papers ↗ · 2026-09-13 Cached

This paper introduces techniques to manage memory peaks in training large Mixture-of-Experts models with long context lengths, including Pipelined LLEP, Ring-DTP, SCO, and OffloadStreamAdamW, which enable fixed GPU working sets and improve throughput up to 10.4x.

0 favorites 0 likes
#training

OpenAI can use all interactions of paid users, even if they opted out of training

Reddit r/LocalLLaMA ↗ · 2026-09-11

The article reveals that OpenAI's terms allow it to use paid users' interactions for training despite opt-out options, due to a loophole with reasoning tokens not being classified as user-owned output.

0 favorites 0 likes
#training

3 ways to prep for your next big race with Search

Google AI Blog ↗ · 2026-09-10 Cached

Google Search's AI features can help runners prepare for races by offering personalized training plans, custom playlists, and gear recommendations.

0 favorites 0 likes
#training

Knowledge Retention & Sharing in DF/IR

Lobsters Hottest ↗ · 2026-09-10 Cached

The article reflects on the historical lack of knowledge sharing and training in Digital Forensics and Incident Response (DF/IR), highlighting personal experiences and efforts to improve processes through automation and documentation.

0 favorites 0 likes
#training

Experimenting with an adaptive memory governor for PyTorch on an 8GB GPU — would love some feedback

Reddit r/LocalLLaMA ↗ · 2026-09-09

The author is experimenting with an adaptive memory governor for PyTorch to prevent CUDA OOM errors on 8GB GPUs, sharing code and seeking community feedback.

0 favorites 0 likes
#training

@LearnOpenCV: 1/20 I trained two segmentation models using GPT-6 Astra Ultra as the orchestrator with no human labels supplied. My pr…

X AI KOLs Timeline ↗ · 2026-09-08 Cached

An experiment where two segmentation models were trained using GPT-6 Astra Ultra as an orchestrator without human labels, with lazy prompts due to time constraints, showing predictions on held-out videos.

0 favorites 0 likes
#training

@AlexGDimakis: 4000 GB200s just arrived in Texas for the Horizon TACC cluster. I'm informed this is the largest academic supercomputer…

X AI KOLs Timeline ↗ · 2026-09-02 Cached

4000 NVIDIA GB200 GPUs arrived in Texas for the Horizon TACC cluster, forming the largest academic supercomputer, with plans to train open AI models.

0 favorites 0 likes
#training

@leeoxiang: A very lightweight TTS implementation, attempted to replicate Audio8's training from scratch using 2000 hours of data. Trained on H200 for less than 10 hours, and the SIM metric already reached 0.72.

X AI KOLs Timeline ↗ · 2026-08-29 Cached

A lightweight TTS implementation, replicating Audio8's training with 2000 hours of data, achieving a SIM metric of 0.72 in under 10 hours of training on H200.

0 favorites 0 likes
#training

@StasBekman: I present to you an Aug 2026 massively revised Machine Learning Engineering open book https://github.com/stas00/ml-engi…

X AI KOLs Following ↗ · 2026-08-26 Cached

This is a massively revised Machine Learning Engineering open book, updated with the latest hardware specifications and examples, providing practical guidance for training and fine-tuning large language models and multi-modal models.

0 favorites 0 likes
#training

Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers

Hugging Face Blog ↗ · 2026-08-26 Cached

This blog post introduces how to train and finetune multi-vector embedding models using the Sentence Transformers library, showcasing its v6.0 update with a new MultiVectorEncoder type and demonstrating superior performance on medical retrieval tasks.

0 favorites 0 likes
#training

Training, learning and inference: unified dynamics of neural systems

arXiv cs.LG ↗ · 2026-08-24 Cached

This paper proposes a unified dynamical framework for training, learning, and inference in neural systems.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback