Tag
This small-scale study evaluates LLM reasoning in legal case forecasting using European Court of Human Rights cases, finding that models produce structurally complete but substantively shallow analyses and that LLM-based evaluators align weakly with human annotators.