Research Feed Page 14
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
Researchers discovered that providing a language model with a prior audit and repair episode causes the model to become more lenient when verifying new mathematical reasoning tasks.
The paper introduces the Semantic Path Compilation system to improve SQL generation reliability by using multi-turn planning and deterministic code-based validation.
The paper introduces Ventor-QTest, an audit framework that detects behavioral shifts in third-party LLM APIs by comparing outputs against trusted benchmarks without needing internal model access.
The paper introduces a diagnostic evaluation suite and a failure taxonomy to systematically identify the root causes of failure in autonomous research agents across the entire scientific lifecycle.
FreeToken enables efficient serving of frontier-scale Mixture of Experts models on consumer hardware by adapting execution strategies to available memory and bandwidth.
TDD-Agent improves repository-level code generation by automating the test driven development process to iteratively refine code and tests.
The paper introduces AnchorBench to evaluate the anchoring effect in large language models across multiple pathways and relevance conditions.
Researchers developed a method using attribute guided genre expansion to train language models on diverse creative formats beyond basic narrative generation.
The paper introduces SPARGen, an instruction-conditioned multimodal generative framework that unifies 3D reconstruction, dense correspondence estimation, and spatial reasoning without task-specific prediction heads.
RecipeNet is a hierarchical transformer model designed to process heterogeneous recipe data with variable schemas and sequential procedural steps.
The paper investigates the disconnect between reasoning behaviors amplified during model training and the actual behaviors that lead to correct answers.
The paper introduces a multi-objective optimization framework to automatically select the best merge parameters for combining specialized neural network models.
The paper introduces a non-parametric approach for multi-modal trajectory prediction that constructs a transition table from historical data to represent uncertainty at route junctions without relying on expensive GPU training or large-scale data.
The paper resolves five deployment bugs and reduces memory overhead to successfully run the Nanbeige4.2-3B model on Apple Silicon.
The paper introduces a hybrid LLM-based framework that automates the generation of SecBPMN2 security annotations from natural-language specifications to improve process model accuracy.
GhostPoint improves 3D object detection by training models to predict the hidden structure of objects that are partially occluded in LiDAR sensor data.
The paper introduces a diffusion transformer method to generate synthetic tabular data by standardizing heterogeneous inputs into a unified statistical format.
MagnifiQ uses a modular patching architecture and LLM-based text prompts to perform efficient high resolution image restoration.
The researchers developed an uncertainty aware classification model to better identify critical high delay cases in business processes that standard regression models often miss.
The paper introduces ReflexVLA, a vision-language-action model architecture that uses future prediction and optimized inference to improve performance in time-sensitive robotics tasks.