Tag
Harvey integrates GPT-6 Astra to enhance AI-driven legal document drafting, providing more context and structured outputs for law firms.
IntLawNER is a new named entity recognition dataset and benchmark for international law, covering gold-annotated sentences from legal texts and evaluating model performance with few-shot improvements.
This paper presents a risk-sensitive evaluation framework for LLM-generated contract clauses, focusing on legal failure modes and quality dimensions to assess risks beyond accuracy or fluency.
Grok 4.7 shows significant improvements in legal work and terminal tasks, outperforming competitors on benchmarks like Harvey Legal Agent Benchmark and Terminal-Bench 4.0, while keeping token prices unchanged.
Vals AI evaluated Grok 4.7, finding it ranks #24 on the Vals Index with a score of 54.2%, down from Grok 4.6, but shows improvements in legal and medical domains.
The paper introduces the PIJ benchmark for evaluating large language models on criminal profiling tasks from incomplete evidence, highlighting performance gaps and biases in inferential reasoning.
A user compares ChatGPT and Gemini, finding Gemini superior in providing accurate legal reasoning and acknowledging logical flaws, while ChatGPT relies on circular arguments and avoids direct answers.
This position paper argues that legal LLM hallucinations should be evaluated as failures of legal warrant rather than factual inaccuracies, proposing a new benchmark framework for assessing legal AI systems.
This paper introduces LexAgentHallu, a hierarchical benchmark for profiling hallucinations in legal AI agents, evaluating multi-step trajectories with fine-grained metrics.
This paper introduces LexIssue, a benchmark for identifying disputed legal issues in Chinese civil litigation, constructed from real cases and expert annotations, and shows that retrieval-augmented generation enhances performance.
The paper introduces OBJECTION, an inference-time pipeline using adversarial lawyer agents to mitigate guilty bias in legal judgment prediction models, demonstrating a significant reduction in false guilty rates and releasing a new 'Natural Innocent' dataset.
This article compares the different strategies of Bloomberg and Thomson Reuters in AI model development, analyzes the shift from training large models from scratch to fine-tuning on open bases, and the impact of this trend on vertical AI applications.
Referent is an AI-native legal practice management platform that combines AI, CRM, and workflow tools to help law firms streamline operations.
Harvey, an OpenAI-backed legal tech firm, has pivoted to using Moonshot AI's Kimi K3 open-weight model to develop its first in-house model, Harvey Tenet, marking a shift towards Chinese open-weight AI systems in Western tech.
Harvey shares strategies for achieving world-class AI research on a limited budget through domain experts guiding synthetic data, establishing evaluation sets, collaborating with labs, and employing model routing.
Introduces ContractScrub, a benchmark for evaluating LLMs on legal contract scrubbing tasks, revealing that current frontier models perform poorly on this domain-specific challenge.
Tenet AI system demonstrates superior performance in corporate law tasks, outperforming Opus5 and Fable across multiple dimensions.
Harvey has post-trained the Kimi K3 model to create Harvey Tenet for long-horizon legal work, achieving improved performance on legal benchmarks and cost-efficiency.
Vector Legal, an AI-native law firm for startups and venture capital, uses LangSmith for agent deployment and fleet management in their legal workflows, having recently raised a seed round.
Thomson Reuters launched the next generation of CoCounsel Legal, an AI-powered product integrating legal research, drafting, and workflows with Westlaw and Practical Law, including the Westlaw Brief Builder for automated drafting while maintaining lawyer control.