100 QFoldAgent: An Autonomous Quantum Optimization Multi-Agent System for Protein Structure Prediction arxiv.org · 3h 4m ago
100 Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy arxiv.org · 3h 4m ago
100 DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs arxiv.org · 3h 4m ago
100 MIITA: Memory-Induced Inference-Time Adaptation for Continual Learning with Small Language Models arxiv.org · 3h 4m ago
100 MedLoCoMo: A Long-Context Multi-Session Medical Dialogue Benchmark for Large Language Models arxiv.org · 3h 4m ago
100 Keyword Matters: Unveiling the Energy Sensitivity of On-Device LLM Prompting arxiv.org · 3h 4m ago
100 SCAIR: Schema-Conditioned Agentic Iterative Reasoning for Enterprise Knowledge Graphs arxiv.org · 3h 4m ago
100 Temporal Context Reinstatement Drives Episodic-Like Order Memory in Long-Context Language Models arxiv.org · 3h 4m ago
100 HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End Optimization arxiv.org · 3h 4m ago
100 Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization arxiv.org · 3h 4m ago
100 The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation arxiv.org · 3h 4m ago
100 Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models arxiv.org · 6d 6h ago
100 FedCC: A Low-Resource Federated Adaptation of Foundation Models for Robust Corpus Callosum localization in Fetal Ultrasound Images arxiv.org · 6d 6h ago
100 BearingNAS: Obtaining In-Sensor Intelligent Fault Diagnosis Systems for Bearings Using a Laptop arxiv.org · 6d 6h ago
100 Dual-domain fused LSTM modeling for efficient time-dependent reliability analysis arxiv.org · 6d 6h ago
100 Reliability Scales Inversely: Bigger Models Compound Mistakes Faster via a Hidden Auto-Regressive Risk Regime arxiv.org · 6d 6h ago
100 Uncertainty Quantification for AI-Driven Crash Simulation Surrogates: A Comparative Study of Monte Carlo Dropout and Deep Ensemble on Open-Source Bumper Beam Benchmark arxiv.org · 6d 6h ago
100 The Information Shadow: Measuring Structural Limits on What Language Models Can Learn arxiv.org · 6d 6h ago
100 Agentic Calibration of Grey-Box Simulation Models: An LLM-Driven Alternative arxiv.org · 6d 6h ago
100 Spatio-Temporal Prediction of Unsteady Airfoil Aerodynamics Using Augmented Graph Neural Ordinary Differential Equations with Exogenous Controls arxiv.org · 6d 6h ago
100 Dynamic Loss Balancing for Joint SOH and RUL Prediction of Lithium-Ion Batteries via a Rotary SOH-Injected Prior Battery Transformer arxiv.org · 6d 6h ago
100 Multi-Timescale Latent-Action DRL for Joint Optimization in Edge-Cloud Networks arxiv.org · 6d 6h ago
100 BatchDAG: LLM-Planned Execution Graphs for Scalable Ad-Hoc Analysis Over Enterprise Data arxiv.org · 6d 6h ago
100 Phionyx: A Deterministic AI Runtime Architecture with Structured State Management and Pre-Response Governance arxiv.org · 6d 6h ago
100 Cross-Dialect Generalization Without Retraining: Benchmarks and Evaluation of Schema-Derived Constrained Decoding for MLIR arxiv.org · 6d 6h ago
100 Semantic Cooperative Games for Contribution Attribution in LLM-Based Multi-Agent Systems arxiv.org · 6d 6h ago
99 “An unprecedented incident.” During a test, an OpenAI model hacked out of its container to reach the internet, then hacked into Hugging Face to steal the test’s answers. reddit.com · 6d 6h ago
91 How does AI find the difference between 2 images? I can understand text response generation and stable diffusion to generate images, but reading and image and finding the differences seems like a whole different beast. reddit.com · 6d 6h ago
57 Nanbeige4.2-3B drops: 3B params claiming to beat 9B/12B models on agentic tasks (atleast according to them) reddit.com · 6d 6h ago
35 Trust but verify doesn’t work when verification of AI processes is difficult theregister.com · 6d 6h ago
34 OpenAI and Hugging Face partner to address security incident during model evalu lesswrong.com · 6d 5h ago
32 AI for Actual Work – a free, self-paced AI training program by Remote.com aiforactualwork.com · 6d 5h ago