large-models

Tag

Cards List
#large-models

@pangyusio: Algorithm engineering is truly a profession where if you don't keep learning, you won't have rice to eat. Algorithm eng…

X AI KOLs Following · 2026-09-01 Cached

This article discusses the necessity for algorithm engineers to continuously learn and adapt to technological shifts, from machine learning to deep learning to large models, with new graduates often leading the adaptation.

0 favorites 0 likes
#large-models

Industry Insights: Niantic Spatial's Big Bet on Large Geospatial Models

Reddit r/artificial · 2026-08-29

Niantic Spatial is focusing on large geospatial models, indicating a strategic shift in AI-driven mapping and spatial technology.

0 favorites 0 likes
#large-models

Large Models for Battery Prognostics and Health Management: A Review and Future Roadmap

arXiv cs.AI · 2026-08-28 Cached

This review paper surveys the application of large models in battery prognostics and health management, addressing long-standing challenges and proposing a roadmap for future research in this domain.

0 favorites 0 likes
#large-models

@Vincent_AINotes: Even if you max out your GPU performance and maintain a steady 100 Tokens per second, running at full load 24 hours a day without interruption, you can only achieve 8.64M Tokens per day. Using this throughput for Agent automation clusters, massive synthetic data generation, or large-scale business analysis is like a drop in the ocean. Local deployment is all about the numbers...

X AI KOLs Following · 2026-08-17 Cached

Local deployment of large models is limited by hardware throughput. Even with GPUs at full capacity, only about 8.64M Tokens can be processed per day, which is insufficient to support Agent automation clusters or large-scale data analysis. Therefore, scaled applications still rely on cloud APIs.

0 favorites 0 likes
#large-models

A Survey of Large Models in Sports

arXiv cs.CL · 2026-08-17 Cached

A comprehensive survey of large models in sports, covering tasks, applications, datasets, and challenges to advance sports intelligence.

0 favorites 0 likes
#large-models

@seclink: Fun fact: Many times, big companies hire in waves. Different companies need different talents in the short term. If you can align with big companies' hiring rhythms, you can often easily land a high salary regardless of education or actual accumulated abilities, surpassing most ordinary people. For example: 1. Recently, Moonshot AI's marketing (especially overseas user growth...

X AI KOLs Following · 2026-08-12 Cached

Fun fact: Domestic big tech companies are currently hiring in batches according to project cycles. Moonshot AI lacks overseas growth, ByteDance focuses on long-term memory and multi-agent tasks, Xiaomi is shifting to B2B for car infotainment/smart home, Tencent Hunyuan lacks pretraining talent, and 3D generation and simulation data talent is scarce with low competition and high salaries.

0 favorites 0 likes
#large-models

@10xmylife: 不发展自己的大模型能行吗?

X AI KOLs Following · 2026-07-26 Cached

Financial Times reports on China's AI talent war, highlighting Moonshot founder Yang Zhilin's ability to retain a top research team amid aggressive poaching by large tech companies.

0 favorites 0 likes
#large-models

What data mix are the labs using to train 10T param models?

Reddit r/singularity · 2026-07-25

Discussion about the data sources labs may use to train 10T parameter models, including synthetic reasoning chains and human-generated traces, amid concerns about hitting the data wall.

0 favorites 0 likes
#large-models

A Survey on the Green Development of Large Models: From Resource-Efficient Architectures to Hardware-Software Co-Design

arXiv cs.LG · 2026-07-13 Cached

This survey comprehensively reviews resource-efficient architectures and hardware-software co-design for green AI, covering efficient model construction, training/deployment strategies, and sustainable hardware, aiming to guide sustainable large model development.

0 favorites 0 likes
#large-models

AI-Model Network: Concept, Current State and Future

arXiv cs.AI · 2026-06-29 Cached

This paper proposes the concept of the world wide AI-Model Network (AI-ModelNet), a novel paradigm for interconnecting, sharing capabilities, and enabling collaborative reasoning among diverse large models. The authors review current single- and multi-model research, present a hierarchical architecture, and validate feasibility through a prototype system and application cases.

0 favorites 0 likes
#large-models

FastMix: Fast Data Mixture Optimization via Gradient Descent

arXiv cs.LG · 2026-06-16 Cached

FastMix is a novel framework that automates data mixture discovery for training large models using a single proxy model and bilevel optimization, achieving state-of-the-art performance with significant efficiency gains.

0 favorites 0 likes
#large-models

@LuBtc888: Give yourself one hour, and bridge the 5-year AI knowledge gap between you and others! DeepMind founder Demis Hassabis's 60-minute talk at Cambridge. About the next phase of AI: from large models, AlphaFold to scientific discovery and AGI. Chinese subtitles added, recommend saving to watch at your leisure.

X AI KOLs Timeline · 2026-06-09 Cached

DeepMind founder Demis Hassabis delivers a 60-minute speech at the University of Cambridge, covering the future development of AI from large models, AlphaFold to scientific discovery and AGI. Video has been added with Chinese subtitles.

0 favorites 0 likes
#large-models

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Hugging Face Daily Papers · 2026-04-15 Cached

Survey introduces the Proxy Compression Hypothesis to explain how RLHF and related methods systematically induce reward hacking, deception, and oversight gaming in large language and multimodal models.

0 favorites 0 likes
← Back to home

Submit Feedback