Research Feed Page 14

Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.

Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Agents / Benchmarks & Evals By Parsa Mazaheri 2026-08-17 2
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency

Researchers discovered that providing a language model with a prior audit and repair episode causes the model to become more lenient when verifying new mathematical reasoning tasks.

Benchmarks & Evals / Reasoning By Yi Ai 2026-08-17
Bounded Semantic Planning and Deterministic Compilation for Reliable Enterprise Text-to-SQL

The paper introduces the Semantic Path Compilation system to improve SQL generation reliability by using multi-turn planning and deterministic code-based validation.

Benchmarks & Evals / Efficiency & Inference By Xiangfan Wu 2026-08-17 12
Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs

The paper introduces Ventor-QTest, an audit framework that detects behavioral shifts in third-party LLM APIs by comparing outputs against trusted benchmarks without needing internal model access.

Agents / Benchmarks & Evals By Yanlin Fei 2026-08-14 19
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks

The paper introduces a diagnostic evaluation suite and a failure taxonomy to systematically identify the root causes of failure in autonomous research agents across the entire scientific lifecycle.

Efficiency & Inference By Shuo Yang 2026-08-17 18
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

FreeToken enables efficient serving of frontier-scale Mixture of Experts models on consumer hardware by adapting execution strategies to available memory and bandwidth.

Agents / Benchmarks & Evals By Hongyue Yu 2026-08-17
TDD-Agent: Test-Driven Reasoning for Code Generation

TDD-Agent improves repository-level code generation by automating the test driven development process to iteratively refine code and tests.

Benchmarks & Evals By Yiderigun Borjigin 2026-08-14
AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs

The paper introduces AnchorBench to evaluate the anchoring effect in large language models across multiple pathways and relevance conditions.

Training & Fine-Tuning By Hwan Chang 2026-08-14
Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion

Researchers developed a method using attribute guided genre expansion to train language models on diverse creative formats beyond basic narrative generation.

Multimodal / Computer Vision By Jinsheng Quan 2026-08-14 1
SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

The paper introduces SPARGen, an instruction-conditioned multimodal generative framework that unifies 3D reconstruction, dense correspondence estimation, and spatial reasoning without task-specific prediction heads.

Training & Fine-Tuning By Pin-Yen Huang 2026-08-14
RecipeNet: A Hierarchical Transformer for Recipe Data

RecipeNet is a hierarchical transformer model designed to process heterogeneous recipe data with variable schemas and sequential procedural steps.

Reasoning / Benchmarks & Evals By Jean de Dieu Nyandwi 2026-08-13 2
Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

The paper investigates the disconnect between reasoning behaviors amplified during model training and the actual behaviors that lead to correct answers.

Training & Fine-Tuning By Utkarsh Agarwal 2026-08-14
Multi-Objective Bayesian Optimization for Model Merging

The paper introduces a multi-objective optimization framework to automatically select the best merge parameters for combining specialized neural network models.

Efficiency & Inference / Multimodal By Michael Fore 2026-08-14
Non-Parametric Spatiotemporal Trajectory Prediction via State-Conditioned Transition Sampling

The paper introduces a non-parametric approach for multi-modal trajectory prediction that constructs a transition table from historical data to represent uncertainty at route junctions without relying on expensive GPU training or large-scale data.

Efficiency & Inference / Agents By John T. Halloran 2026-08-14
Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead

The paper resolves five deployment bugs and reduces memory overhead to successfully run the Nanbeige4.2-3B model on Apple Silicon.

Safety & Alignment / Benchmarks & Evals By Md Kamrul Islam 2026-08-14
A Hybrid LLM-Based Framework for Automated Security Annotation Generation in Business Process Models

The paper introduces a hybrid LLM-based framework that automates the generation of SecBPMN2 security annotations from natural-language specifications to improve process model accuracy.

Computer Vision / Training & Fine-Tuning By Mohamed Abdelsamad 2026-08-14
GhostPoint: Self-Supervised Representation Learning by Hallucinating Occluded LiDAR Structure

GhostPoint improves 3D object detection by training models to predict the hidden structure of objects that are partially occluded in LiDAR sensor data.

Benchmarks & Evals By Hao Yan 2026-08-14
Generating Benchmark Health Data Using a Tabular Diffusion Transformer

The paper introduces a diffusion transformer method to generate synthetic tabular data by standardizing heterogeneous inputs into a unified statistical format.

Computer Vision / Efficiency & Inference By Mahesh Reddy 2026-08-14
MagnifiQ: Patch-aware Text Guided Progressive Upscaling for High-Resolution Image Restoration

MagnifiQ uses a modular patching architecture and LLM-based text prompts to perform efficient high resolution image restoration.

Benchmarks & Evals By Keyvan Amiri Elyasi 2026-08-14
Mind the Long Tail: Understanding the Difficulty of Delay Detection in Business Processes

The researchers developed an uncertainty aware classification model to better identify critical high delay cases in business processes that standard regression models often miss.

Robotics / Efficiency & Inference By Yuxuan Chen 2026-08-14
Reflex: Enabling Fast and Predictive Vision-Language-Action Models for Reaction-Critical Manipulation

The paper introduces ReflexVLA, a vision-language-action model architecture that uses future prediction and optimized inference to improve performance in time-sensitive robotics tasks.