partial-evaluation

Tag

Cards List
#partial-evaluation

A First Futamura Projection

Lobsters Hottest · 2d ago Cached

The blog post explains how to implement the first Futamura projection by specializing an interpreter to a program, using a Brainfuck-to-Carp compiler as a practical example.

0 favorites 0 likes
#partial-evaluation

How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks

arXiv cs.AI · 2026-07-15 Cached

This paper analyzes how many tasks are needed in partial evaluations of LLM agent benchmarks to reach the same pairwise conclusions as full benchmarks. It finds that required task fractions vary sharply across benchmarks and suggests reporting standards for partial evaluations.

0 favorites 0 likes
← Back to home

Submit Feedback