Research Feed Page 40
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
Tytan uses neurosymbolic AI to automatically build semantic schemas from raw relational databases by combining LLM-driven inference with deterministic verification.
The paper presents a framework using generative AI and JSON schemas to automate the extraction and semantic evaluation of complex hierarchical data from health technology assessment documents.
VideoArgus introduces an agentic evaluation framework that uses dynamic, instance specific rubrics and specialized tools to provide accurate, diagnostic feedback on video generation and editing tasks.
StreamArena is a new benchmark for evaluating agentic streaming video understanding, and StreamMind is a two-tier architecture designed to optimize latency and performance for these long-horizon tasks.
LILAC is a neural speech codec designed to be idempotent, ensuring that repeated cycles of decoding and re-encoding audio do not degrade signal quality or diverge in token streams.
The paper introduces a framework to evolve agent skills as an interconnected global system rather than isolated updates to improve performance and generalizability.
The paper introduces BaKron, a new quantization method that improves efficiency for two-sided Kronecker-factored Hessian approximations in neural networks.
The paper introduces a method called SO-OPF to decompose vision encoder representations into distinct support and operation factors to evaluate how well models generalize to new visual configurations.
Argus is a persistent runtime system that improves research agent performance by evolving operational state and project objectives alongside human guidance.
Researchers improved language model learning efficiency by pre-pretraining a Transformer backbone on formal logic derivation sequences before standard language training.
The authors introduce GDPevo, a benchmark and automated pipeline designed to evaluate and improve how agents evolve their performance on complex enterprise workflows.
The paper demonstrates that using LoRA adapters can rewrite AI-generated text to match a specific user's writing style without needing explicit style instructions.
The paper introduces a method that uses grounded transformation chains to supervise intermediate reasoning steps for grid-based visual puzzles.
The authors introduce AgentHPOBench to evaluate how effectively LLM agents perform sequential hyperparameter optimization across thirty machine learning tasks.
The researchers demonstrate that robots can learn effective manipulation policies using only high-fidelity handheld video demonstrations instead of expensive real-robot teleoperation data.
OmniDelta optimizes token compression in audio-video large language models by dynamically allocating processing budgets based on task-specific relevance.
The paper introduces a stacked architecture for quantum algorithms that allows users to adjust the trade-off between the ease of training a model and its resistance to being simulated by classical computers.
ID-V2V is a generative framework that uses multi-stream control signals to restyle videos while maintaining strict subject identity and performance.
The paper demonstrates that using a three-stage multi-agent pipeline instead of a single model call significantly changes how models align with specific target objectives.
The researchers developed Diff-Logic, a method for running EEG classification on edge devices by replacing heavy floating-point arithmetic with sparse Boolean circuits.