Literature Review: MELTing point: Mobile Evaluation of Language Transformers | Bnechmarking LLMs on Phones
Summary
This literature review introduces MELT, a benchmark for evaluating large language models (LLMs) on mobile devices, highlighting the challenges and opportunities of running transformers on phones.
Similar Articles
Literature Review: LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load | Bnechmarking LLMs on Phones [R]
This literature review analyzes performance efficiency trade-offs for LLM inference on mobile devices, NPUs, and GPUs under sustained load, focusing on benchmarking LLMs on phones.
Transformer-Based Language Models Across Domain Verticals: Architectures, Applications and Critical Assessment
A comprehensive survey of transformer-based language models covering architectures, applications across domain verticals (healthcare, finance, legal, etc.), and critical assessment of trade-offs including compute cost, alignment, and data provenance.
(Towards) Scalable Reliable Automated Evaluation with Large Language Models
This paper proposes a scalable, domain-agnostic framework for automated LLM evaluation that uses pairwise comparisons by multiple LLMs and an Elo rating system to approximate expert judgments, reducing the need for human intervention.
Benchmarking LLMs
A study or report on benchmarking large language models, likely comparing performance across various tasks.
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
MetroLLM-Bench is a 955-case benchmark for evaluating language models as transit kiosk policy layers, showing that small fine-tuned models can match larger models on structured tool-use and fare-quoting tasks.