Research Feed Page 26
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
Spark-to-Paper integrates paper generation into existing coding assistants as a composable workflow that verifies experimental results and minimizes hallucination.
QV-PIC is a query-aware caching framework that recovers lost textual details in visual inputs to improve RAG latency and quality.
The paper introduces a method to group control transitions in LLM agents to improve GPU execution efficiency and reduce latency by avoiding host round trips.
An agentic workflow successfully modernized tens of thousands of lines of legacy Fortran code by using specialized agent roles, version-controlled specifications, and exact verification.
VAKRA evaluates how well AI agents perform complex multi-step reasoning by combining structured API calls with document retrieval.
The paper introduces SCOPE-Router, a cost-aware system that assigns tasks to the most suitable vision-language models for execution-oriented workflows.
Researchers developed a method where a strong builder model automatically generates and refines scaffolding logic to improve the performance of smaller language models on reasoning tasks.
The paper introduces a dedicated, accurate standalone visibility estimation method for hand keypoints that outperforms existing auxiliary approaches.
The paper introduces a system of preventive and evidential layers to secure autonomous agents through runtime contracts rather than relying on training-time model alignment.
ToolHazard provides a scalable framework to automatically synthesize stateful environments and generate adversarial tasks for evaluating and aligning LLM-based agents.
The paper introduces FACT, a causal world model that improves robotic action planning by explicitly learning from both successful and failed outcomes.
The researchers developed Sekai2, a large-scale video dataset designed to provide the temporal continuity and camera data necessary for training interactive world models.
The researchers developed a metric called VIScore to diagnose how well latent world models translate their internal representations into effective planning for robotics tasks.
The paper introduces Cultivar, a new evaluation framework that detects data contamination and measures how well translation models handle locale-specific cultural nuances.
The researchers constructed a Burmese medical speech corpus and fine-tuned Whisper models to improve automatic speech recognition for clinical dialogues.
The researchers developed a World-Action Model that leverages action-free video pretraining to significantly boost the performance of surgical robots when labeled demonstration data is limited.
The researchers developed the HUI360 dataset and baseline models to predict when humans will physically interact with a mobile robot.
CausalSplat enables 3D Gaussian Splatting systems to interpret complex user instructions by mapping visual data to a structured scene graph.
The researchers developed a reinforcement learning approach that uses verifiable temporal grounding to improve the accuracy of detecting AI-generated video forgeries.
The paper introduces a framework that improves patent matching accuracy by using an LLM-driven process to mine technical entities and construct hierarchical ontologies for enhanced query retrieval.