Research Feed Page 28
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The authors introduce Workflow Cards, a structured metadata format that improves how provenance data is captured and interpreted by large language models.
The researchers investigated whether English safety alignments carry over to low-resource languages by testing models on the newly compiled LoDNA dataset.
The paper introduces a cost-aware routing strategy that selects between direct multilingual processing and translation-based English classification to optimize performance for weaker languages.
The paper introduces a synthetic generation and verification framework called V-FiLLM to benchmark and improve large language model reasoning over structured financial data.
The paper introduces a diagnostic and remediation framework called VISOR that identifies and corrects attribute hallucination errors in vision-language models by distinguishing between language-layer biases and visual representation failures.
The researchers developed a method to quantify how different language models behave and evolve by measuring distances between their responses to a shared set of 10,000 prompts.
PEAK uses sparse autoencoders to precisely identify and suppress target concepts in diffusion models while maintaining overall generation quality.
SafeCA is a defensive framework that regulates cross-attention mechanisms in text-to-video generative models to prevent the output of harmful or inappropriate content.
Researchers identified specific pre-training documents that cause emergent misalignment in language models and demonstrated that synthetic instruction tuning exacerbates this behavior.
The paper introduces a composition-based framework that allows robots to dynamically discover and integrate new software and hardware payloads for task execution at runtime.
XCoT-VLA replaces verbose natural-language reasoning with compact, executable tokens to improve driving performance and inference efficiency.
SCOUT is a system that identifies and localizes latent hardware or communication failures during distributed large language model pre-training by comparing the behavior of identical parallel processing units.
ThinkRetrieve improves large model reasoning by dynamically injecting relevant, solved examples into the reasoning process at each step.
ReRound uses a learned diffusion-based approach to resolve midpoint ambiguity during model quantization, resulting in higher accuracy for compressed LLMs without requiring calibration data.
The paper introduces a scheduling method that increases rollout throughput and reduces iteration time by managing how heterogeneous reinforcement learning workloads share KV-cache capacity.
The researchers developed AdvFD, a new training method that prevents image generators from exploiting static metrics to inflate their performance scores without actually improving visual quality.
The paper introduces Power Law Graph Attention as a flexible, learned alternative to the standard fixed-operator attention used in modern transformer models.
The paper demonstrates that using templated prompts creates structural artifacts that bias LLM political stance measurements, whereas LLM-generated prompts produce more realistic and neutral results.
The paper introduces Latent-to-4D, a framework that uses a shared latent space to enable a single geometry-supervised 4D model to work across multiple compatible video diffusion transformers.
This study exposes how undocumented configuration differences and inconsistent evaluation protocols significantly alter the reported performance of the LeWorldModel agent.