Research Feed Page 25
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The paper introduces Generation as Auxiliary Supervision, a method that uses a decoupled generation branch during training to improve Multimodal Large Language Model visual understanding while adding zero overhead at inference time.
The paper introduces a statistical method to control error rates in large language models by dynamically selecting which outputs to trust.
VICBench is a new multi-language benchmark containing 100 verified vulnerability-inducing commits that helps evaluate automated security detection tools.
The paper examines adoption and usage patterns of ChatGPT Enterprise by analyzing internal message data and firm-level financial information.
LoSA accelerates video diffusion models by identifying and caching the most important attention blocks to reduce the overhead of quadratic computation in long 3D token sequences.
GUIDE is an automated framework that parses, validates, and converts heterogeneous enterprise documents into structured, deployment-ready artifacts.
The researchers introduced Claim Level Reliability Assessment, a method that improves language model reasoning accuracy and efficiency by verifying individual logical steps rather than relying on final trace results.
The paper introduces Context-Calibrated DPO to force models to better utilize contextual information and reduce object hallucinations.
The paper introduces Graph-Structured Rubrics to replace flat evaluation criteria with a deterministic, auditable directed acyclic graph that improves scoring accuracy.
The paper introduces SAG, a retrieval system that improves RAG performance by using SQL joins to dynamically discover cross-document associations through event-based hyperedges.
The SHAPER method improves agent performance in new environments by evolving textual skills and harnesses while keeping the underlying model parameters frozen.
The authors introduce RealisticTritonBench to evaluate LLM performance in generating production-grade Triton kernels for real-world AI frameworks.
Researchers developed VITA, a domain-specific retrieval-augmented generation system that achieves superior performance on clinical tasks by prioritizing curated local data over generic model scale.
StateFlow introduces a framework for generating and interactively editing persistent 3D world states to solve the controllability and consistency issues found in one-shot video synthesis.
The paper introduces SoftWater, a class-aware rate allocation algorithm for quantisation that reduces memory usage in the softmax output layer of Large Language Models.
The paper demonstrates that model rankings shift significantly based on the token generation budget allowed during inference, challenging the reliability of standard static evaluation benchmarks.
The researchers found that training language models on long documents makes them rely more on provided text and less on their own internal knowledge, leading to worse performance when that context is missing.
The study demonstrates a framework called Genesis that enables autonomous agents to build complex software by using persistent recursive states and iterative validation to evolve codebases from empty repositories.
The paper introduces a method called Convergent Detour Hijacking that steers LLM agents into unnecessarily costly execution paths while preserving the final task output.
Gambit improves the accuracy of large reasoning models by dynamically reallocating computational resources to promising branches during inference.