Research Feed Page 21
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
Researchers developed an algebraic method to identify which regular languages allow transformers to generalize to sequence lengths beyond their training data.
SteerBench-Work provides a structured, incident-anchored benchmark to evaluate whether AI agents correctly decide to execute or pause tasks for human review.
FIRE-VLA improves autonomous driving models by using a distillation method to correct consistent failures where reinforcement learning signals are insufficient.
The paper models how engineers should shift resources between training-time character shaping and inference-time rule enforcement as system deployment scale increases.
The paper introduces drift-aware retraining strategies to maintain malware classification accuracy while minimizing the number of model updates compared to periodic retraining.
The paper introduces a framework to evaluate AI research agents by decomposing their workflow into specific capabilities and measuring process reliability rather than just final outputs.
The paper introduces Princigram, a framework that uses structured physics-based constraints to improve the fidelity of scientific images generated by multimodal models.
TopoIntent translates natural language requirements into validated, compliant network topologies using a multi stage pipeline of intent analysis, template retrieval, and iterative repair.
SynWeaver improves web agent accuracy by co-synthesizing website-specific tasks and execution trajectories to overcome the lack of supervision on unseen websites.
The paper introduces a method to prevent document multimodal large language models from leaking correlated sensitive fields when given abnormal inputs.
CoverPrune improves 3D vision model efficiency by using optimal transport to intelligently prune visual tokens while preserving essential scene coverage.
The paper introduces a distillation method that aligns teacher supervision with causal inference to resolve context mismatches in video generation models.
OmniScientist is an autonomous agent system that processes raw multimodal scientific data to generate hypotheses, execute experiments, and write manuscripts.
ContactGuard uses a predictive world model to detect and abort robot contact failures before the physical interaction occurs.
The paper introduces H2R-Bench to measure how effectively video world models translate human hand manipulation into robot-centric motor task videos.
The paper introduces a fine-tuning method to prevent models from being tricked by malicious prompt wrappers that bypass safety filters or cause over-refusal of benign tasks.
The researchers evaluated whether large language models can intentionally provide less specific answers when they encounter entities outside their training knowledge to avoid hallucinations.
SPADE lowers inference costs and latency for large language models by using an edge-based draft model to generate token sequences that are verified by a cloud-based model.
AaLLM is an end-to-end framework that uses a chain of LLM agents to automatically generate and size analog circuits from user specifications.
The paper introduces a contract-based verification system that prevents LLM-powered proof repair tools from making unauthorized edits to protected code in the Isabelle assistant.