spanish

Tag

Cards List
#spanish

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use

arXiv cs.CL · 2026-08-11 Cached

This paper introduces VectraYX-Vision-1B, a sub-2B Spanish/LATAM cybersecurity vision-language model, but reports a negative visual-grounding result, raising architectural questions about NoPE layers and releasing code, benchmarks, and checkpoints.

0 favorites 0 likes
#spanish

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use

Hugging Face Daily Papers · 2026-08-09 Cached

Presents VectraYX-Vision-1B, a sub-2B Spanish/LATAM cybersecurity vision-language model coupling a frozen SigLIP encoder with a Spanish decoder via an MLP, yet reports near-zero visual grounding despite functional pipelines, with open-source weights and remediation plans.

0 favorites 0 likes
#spanish

Probing Character-level Transformers for the Spanish L-shaped Morphome

arXiv cs.CL · 2026-08-05 Cached

This paper probes character-level transformers to investigate whether they encode the Spanish L-shaped morphome, an irregular morphological pattern, as an abstract class or just surface alternations. The authors find that the encoding is item-specific and localized, but does not generalize like human learners.

0 favorites 0 likes
#spanish

MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams

arXiv cs.CL · 2026-07-22 Cached

Introduces MIRA-Ev, a clinical argument mining benchmark built on Spanish MIR licensing-exam cases, annotated with span-level premises, claims, and support/attack relations, available in Spanish, English, and Basque. It provides a three-tier task hierarchy for evaluating evidence sentence retrieval, argumentative component extraction, and relation classification, addressing the limitations of multiple-choice QA benchmarks in clinical NLP.

0 favorites 0 likes
#spanish

ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions

arXiv cs.CL · 2026-07-21 Cached

Introduces ESCUCHA, the first Spanish speech understanding benchmark for evaluating large audio language models across heterogeneous acoustic conditions and reasoning abilities, comprising 1,000 curated questions from diverse real-world sources.

0 favorites 0 likes
#spanish

La IA está premiando el volumen y enterrando la innovación

Reddit r/artificial · 2026-07-13

Artículo que critica cómo la inteligencia artificial está premiando el volumen de datos y producción sobre la innovación real, sugiriendo un desequilibrio en el campo.

0 favorites 0 likes
#spanish

Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration

arXiv cs.CL · 2026-07-10 Cached

Introduces a cost-efficient human-LLM collaborative annotation framework to construct EspanStereo, a Spanish-language stereotype dataset covering multiple Spanish-speaking countries, enabling more culturally grounded bias evaluation in LLMs.

0 favorites 0 likes
#spanish

S-DiverSe: Spanish Diverse Speech

arXiv cs.CL · 2026-07-07 Cached

S-DiverSe is a 3.2-hour corpus of Spanish speech from 22 speakers with neurological conditions (ALS, Parkinson's, stroke), designed to support ASR evaluation for pathological speech. Baseline experiments show heuristic post-processing outperforms fine-tuning for this domain.

0 favorites 0 likes
#spanish

HULAT2 at MER-TRANS 2026: Governed Multi-Agent Simplification for Spanish Easy-to-Read Generation

arXiv cs.CL · 2026-07-03 Cached

This paper describes HULAT2-UC3M's participation in the MER-TRANS 2026 shared task on Spanish Easy-to-Read generation, using a governed multi-agent workflow with LangGraph and Gemini/RigoChat models, achieving best SARI of 44.05.

0 favorites 0 likes
#spanish

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use

arXiv cs.CL · 2026-05-15 Cached

Presents VectraYX-Nano, a 42M-parameter decoder-only language model trained from scratch in Spanish for cybersecurity, featuring curriculum learning, native tool invocation via MCP, and a 170M-token corpus. Empirical findings reveal a loss-versus-register inversion and corpus-density artifacts for tool-use capability.

0 favorites 0 likes
← Back to home

Submit Feedback