@askalphaxiv: "Atomistic Language Models Understand and Generate Materials" Most materials AI still treats crystals and language sepa…

X AI KOLs Timeline Papers

Summary

This paper introduces an atomistic language model that integrates a 3D atom encoder, Qwen LLM, and diffusion crystal generator to natively handle multimodal materials data, achieving state-of-the-art crystal structure prediction and de novo generation.

"Atomistic Language Models Understand and Generate Materials" Most materials AI still treats crystals and language separately, either turning atoms into lossy text formats or making LLMs call atomistic tools. This paper makes materials natively multimodal by connecting a 3D atom encoder, Qwen LLM, and diffusion crystal generator through continuous latent projectors. This model can read atomic coordinates, predict properties, edit crystals from text, and generate stable new materials, with SoTA crystal structure prediction and strong de novo generation.
Original Article
View Cached Full Text

Cached at: 06/24/26, 06:28 PM

“Atomistic Language Models Understand and Generate Materials”

Most materials AI still treats crystals and language separately, either turning atoms into lossy text formats or making LLMs call atomistic tools.

This paper makes materials natively multimodal by connecting a 3D atom encoder, Qwen LLM, and diffusion crystal generator through continuous latent projectors.

This model can read atomic coordinates, predict properties, edit crystals from text, and generate stable new materials, with SoTA crystal structure prediction and strong de novo generation.

Similar Articles

LapidaryEngine: Fully Conversational Crystal Generation

arXiv cs.LG

LapidaryEngine is a new AI model that enables fully conversational generation of crystal materials from free-form natural language, using a pivot representation for bidirectional translation and iterative refinement. It outperforms existing text-to-crystal systems by allowing intuitive, dialogue-like interaction.

Miller-Index-Based Latent Crystallographic Fracture Plane Reasoning with Vision-Language Models

arXiv cs.LG

This paper investigates whether multimodal large language models (MLLMs) can leverage Miller indices as a latent representation to reason about crystallographic fracture geometry from visual inputs, evaluating their ability to infer physically valid plane hypotheses and determine when such representation is applicable across materials like ceramics, glass, metals, and concrete.