Research Feed Page 9

Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.

Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Safety & Alignment By Matthew Faucher 2026-08-21
TRACE-C: Rank-Calibrated Relational Anomaly Detection for Multi-Stream Operational Telemetry

The paper introduces a rank-calibrated detector called TRACE-C designed to identify anomalies in complex electricity system telemetry.

Efficiency & Inference / Benchmarks & Evals By Mieszko Komisarczyk 2026-08-21
Tydra: An Efficient Hybrid Model for Tabular Data

Tydra combines transformer and state-space architectures to achieve faster inference on tabular data than the existing TabPFN foundation model.

Agents / Reinforcement Learning By Jiakai Tang 2026-08-21 4
Towards Faithful Simulation of Human Shopping Behavior

The paper introduces a GUI-grounded simulation agent that uses pixel-level perception and reinforcement learning to generate authentic, multi-turn e-commerce shopping trajectories.

Benchmarks & Evals / Safety & Alignment By Junseok Kim 2026-08-21
Personalized Privacy Control in LLMs via Attention Head Intervention

The paper introduces a method to improve privacy policy adherence in LLMs by intervening on specific attention heads to align model outputs with user-defined privacy preferences.

Multimodal / Benchmarks & Evals By Enjun Du 2026-08-21 8
EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking

EviRank improves image retrieval accuracy by replacing unstructured reasoning with a structured, criteria-based verification framework.

Computer Vision / Efficiency & Inference By Yunze Tong 2026-08-21 11
InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter

InfinityEdit uses a lightweight adapter to enable consistent, long-term video editing for continuous data streams.

Agents / Benchmarks & Evals By Adriana Watson 2026-08-21
From Regulation to Implementation: A Critical Evaluation of LLM-Assisted Regulatory Compliance in Industry

This research evaluates how effectively various language models automate the generation of compliance documentation like digital product passports and data protection assessments.

Multimodal / Benchmarks & Evals By Yibo Hu 2026-08-21 7
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

TLive-Omni is a multimodal model designed to process and understand the complex mix of speech, text, and video signals found in e-commerce live streams.

Efficiency & Inference By Peiqi Yu 2026-08-21
COEC: Calibrated Orthogonal-Equivalence Compensation for Structured Pruning of Large Language Models

The COEC method improves the accuracy of pruned large language models by using a specialized, two-sided rotation technique that avoids the pitfalls of direct weight refitting.

Agents / Benchmarks & Evals By Lekang Jiang 2026-08-21
Benchmarking Patent Drafting from Inventor-Style Disclosures

The authors introduce a multi-agent system called Patent-MAF that processes raw invention disclosures into formal patent specifications and claims.

Training & Fine-Tuning / Efficiency & Inference By Zeyun Zhong 2026-08-21
Rethinking Expressivity and Efficiency in Test-Time Training

The paper introduces E2-TTT, a new method for Test-Time Training that uses chunk-wise updates to achieve higher performance while maintaining computational efficiency.

Multimodal / Benchmarks & Evals By Xianyun Sun 2026-08-21
OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs

OmniAssistBench is a new evaluation framework designed to measure how well multimodal AI models handle complex, multi-turn interactions in real-time video scenarios.

Agents / Benchmarks & Evals By Oleg Grynets 2026-08-21
Specification Portability Across LLM Development Agents: Cross-Agent Compatibility in Specification-Driven Software Migration

The researchers investigated whether software migration specifications are interchangeable across different LLM development agents, finding that they are not agent-neutral artifacts.

Agents / Safety & Alignment By Yingzhe Tong 2026-08-21
AID-Guard: Stateful Authorization for Delegated Agent Effects

AID-Guard ensures that AI agent decisions result in exactly one provider effect by binding user intent to durable, stateful authorization protocols.

Robotics / Efficiency & Inference By Zhuoyuan Li 2026-08-21
Just Noticeable Difference Modeling for Token Compression in Vision-Language-Action Models

The authors introduce a method to compress token data in vision-language-action models by identifying and prioritizing information that has the least impact on physical robot movements.

Agents / Safety & Alignment By Arulnidhi Karunanidhi 2026-08-21
Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking

The paper demonstrates that existing content screening and provenance ranking methods fail to reliably defend agent memory systems from adversarial data injection.

Safety & Alignment By Zhibo Zhang 2026-08-21
RARE: Decoupling Representation Steering from Expert Routing in Mixture-of-Experts Language Models

The paper introduces RARE, a method to steer Mixture of Experts models by decoupling control interventions from the model router mechanism to maintain performance and reliability.

Reasoning / Benchmarks & Evals By Xuanyu Meng 2026-08-21
EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering

EnSI-RAG improves long-document question answering by indexing documents based on structured entity relationships rather than simple text chunks.

Safety & Alignment / Benchmarks & Evals By Balkrishna Giri 2026-08-21
Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems

The paper introduces an agent that evaluates Retrieval-Augmented Generation outputs by combining document screening and factual verification to block poisoned data and instruction injection.

Agents / Reasoning By Jason Hickey 2026-08-21
AI with Authority, from Application to Silicon

A single researcher successfully used AI agents to develop a complete system from application code to silicon tapeout in five weeks.