Research Feed Page 30
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The paper introduces VibeLifeBench, a framework for evaluating agent performance in long-term, multi-week scenarios that involve state persistence and silent environment changes.
Researchers developed a method to isolate and steer specific internal features within multimodal models to improve performance on spatial and OCR tasks without retraining.
The paper introduces a hybrid system that splits optimization tasks into an LLM-driven structural design phase and a traditional numerical parameter tuning phase to prevent structural failures.
The paper demonstrates that LLM-driven search processes often optimize for specific benchmark configurations rather than generalizing to unseen settings.
AnyCamVLA improves robot task performance in new camera environments by synthesizing training-viewpoint images in real-time before processing them with a pre-trained policy.
The paper introduces a benchmark called PragMatch to test if large vision-language models can distinguish between genuine pragmatic sarcasm and simple image-text mismatches.
The paper introduces a new benchmark to evaluate how well large language models can resolve specific versions of legal documents that change over time.
ColluSkill exploits the isolation of current agent skill scanners by decomposing malicious intents into interdependent sub-payloads that bypass detection as individual, locally benign components.
Researchers developed a new, systematic evaluation suite to help Dutch government agencies select and deploy large language models based on transparency, honesty, and operational efficiency.
Researchers successfully used a single causal language model to perform stock forecasting and portfolio allocation by treating financial data as tokens rather than using traditional task-specific numerical heads.
DUET improves video generation quality and diversity in a two-step process by combining two distinct distillation expert models.
LookAgain introduces a multi-turn refinement process that allows GUI agents to visually verify and adjust their coordinate predictions, significantly boosting accuracy on complex interfaces.
The paper evaluates how different governance structures, such as constitutional prompts and provenance-aware guards, impact safety compliance in multi-agent delegation workflows.
KGCaRe enhances complex conditional reasoning in language models by combining structured knowledge graph lookups with traditional document retrieval.
XPolicyLab introduces a unified standard and client/server architecture to resolve the fragmentation in robot policy integration and evaluation.
VectraYX-Vision-1B is a specialized vision language model designed for offline cybersecurity reasoning and native tool invocation in Spanish and Latin American contexts.
Researchers discovered that tool use in vision language models relies on structured text scaffolding rather than the actual images returned by those tools.
The paper introduces a joint loss optimization technique that fine-tunes LLMs to improve task accuracy while simultaneously minimizing inference carbon emissions.
The researchers investigated how to balance a model's ability to provide concise answers with its capacity for long-form mathematical reasoning by testing different training schedules and data ratios.
ReMEMBER is a method that improves streaming dialogue summarization by retrieving targeted evidence for unresolved context gaps in long conversations.