Research Feed Page 21

Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.

Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Benchmarks & Evals By Andy Yang 2026-08-13
Algebraic Decomposition Theory for Transformer Length Generalization

Researchers developed an algebraic method to identify which regular languages allow transformers to generalize to sequence lengths beyond their training data.

Agents / Benchmarks & Evals By Oguz Serdar 2026-08-12
SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries

SteerBench-Work provides a structured, incident-anchored benchmark to evaluate whether AI agents correctly decide to execute or pause tasks for human review.

Reinforcement Learning / Computer Vision By Hao Dou 2026-08-13
FIRE-VLA: Failure-Informed Self-Evolution for Vision-Language-Action Models in Autonomous Driving

FIRE-VLA improves autonomous driving models by using a distillation method to correct consistent failures where reinforcement learning signals are insufficient.

Safety & Alignment / Reinforcement Learning By Satoshi Takahashi 2026-08-13
Rules or Character? Scaling Laws for AI Safety Design

The paper models how engineers should shift resources between training-time character shaping and inference-time rule enforcement as system deployment scale increases.

Training & Fine-Tuning / Efficiency & Inference By Christofer Washington Berruz Chungata 2026-08-13
Concept Drift Detection and Adaptive Retraining of Malware Classification Models

The paper introduces drift-aware retraining strategies to maintain malware classification accuracy while minimizing the number of model updates compared to periodic retraining.

Agents / Benchmarks & Evals By Yiwei Li 2026-08-13
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

The paper introduces a framework to evaluate AI research agents by decomposing their workflow into specific capabilities and measuring process reliability rather than just final outputs.

Multimodal / Benchmarks & Evals By Minghui Zhang 2026-08-13 1
Towards Physics-Faithful Generation of Scientific Diagrams

The paper introduces Princigram, a framework that uses structured physics-based constraints to improve the fidelity of scientific images generated by multimodal models.

Agents / Benchmarks & Evals By Xiaokang Qu 2026-08-13
TopoIntent: Compiling Security Intent into Executable, Compliance-Checked Network Topologies

TopoIntent translates natural language requirements into validated, compliant network topologies using a multi stage pipeline of intent analysis, template retrieval, and iterative repair.

Agents / Benchmarks & Evals By Ruitao Wang 2026-08-12
SynWeaver: Website-Prior Task and Trajectory Co-Synthesis for Web Agents

SynWeaver improves web agent accuracy by co-synthesizing website-specific tasks and execution trajectories to overcome the lack of supervision on unseen websites.

Safety & Alignment / Multimodal By Beining Xu 2026-08-13
Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs

The paper introduces a method to prevent document multimodal large language models from leaking correlated sensitive fields when given abnormal inputs.

Multimodal / Efficiency & Inference By Peng Ling 2026-08-13
CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport

CoverPrune improves 3D vision model efficiency by using optimal transport to intelligently prune visual tokens while preserving essential scene coverage.

Computer Vision / Efficiency & Inference By Hmrishav Bandyopadhyay 2026-08-13 1
Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation

The paper introduces a distillation method that aligns teacher supervision with causal inference to resolve context mismatches in video generation models.

Agents / Multimodal By Bobo Li 2026-08-13 2
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

OmniScientist is an autonomous agent system that processes raw multimodal scientific data to generate hypotheses, execute experiments, and write manuscripts.

Robotics / Safety & Alignment By Gehan Zheng 2026-08-13
ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models

ContactGuard uses a predictive world model to detect and abort robot contact failures before the physical interaction occurs.

Robotics / Benchmarks & Evals By Dingyi Rong 2026-08-13 6
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

The paper introduces H2R-Bench to measure how effectively video world models translate human hand manipulation into robot-centric motor task videos.

Training & Fine-Tuning / Safety & Alignment By Ping Wu 2026-08-13
Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety

The paper introduces a fine-tuning method to prevent models from being tricked by malicious prompt wrappers that bypass safety filters or cause over-refusal of benign tasks.

Benchmarks & Evals / Safety & Alignment By Dananjay Srinivas 2026-08-13
Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity

The researchers evaluated whether large language models can intentionally provide less specific answers when they encounter entities outside their training knowledge to avoid hallucinations.

Efficiency & Inference / Benchmarks & Evals By Divya Jyoti Bajpai 2026-08-13
SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference

SPADE lowers inference costs and latency for large language models by using an edge-based draft model to generate token sequences that are verified by a cloud-based model.

Agents / Efficiency & Inference By Mohammed Ayman Habib 2026-08-13
AaLLM: An End-to-End Analog Circuit Design Framework from Topology Generation to Sizing Using Large Language Models

AaLLM is an end-to-end framework that uses a chain of LLM agents to automatically generate and size analog circuits from user specifications.

Agents / Safety & Alignment By Jim Woodcock 2026-08-13
CAPRI: Contract-Aware Proof Repair for Isabelle

The paper introduces a contract-based verification system that prevents LLM-powered proof repair tools from making unauthorized edits to protected code in the Isabelle assistant.