100 QFoldAgent: An Autonomous Quantum Optimization Multi-Agent System for Protein Structure Prediction arxiv.org · 1h 16m ago
100 Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy arxiv.org · 1h 16m ago
100 DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs arxiv.org · 1h 16m ago
100 MIITA: Memory-Induced Inference-Time Adaptation for Continual Learning with Small Language Models arxiv.org · 1h 16m ago
100 MedLoCoMo: A Long-Context Multi-Session Medical Dialogue Benchmark for Large Language Models arxiv.org · 1h 16m ago
100 Keyword Matters: Unveiling the Energy Sensitivity of On-Device LLM Prompting arxiv.org · 1h 16m ago
100 SCAIR: Schema-Conditioned Agentic Iterative Reasoning for Enterprise Knowledge Graphs arxiv.org · 1h 16m ago
100 Temporal Context Reinstatement Drives Episodic-Like Order Memory in Long-Context Language Models arxiv.org · 1h 16m ago
100 HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End Optimization arxiv.org · 1h 16m ago
100 Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization arxiv.org · 1h 16m ago
100 The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation arxiv.org · 1h 16m ago
100 ParBench: A Benchmark for Reliable Evaluation of LLM Parallel Code Translation arxiv.org · 1h 16m ago
100 Lexical discovery in unknown environments orchestrated by Large Language Models arxiv.org · 1h 16m ago
100 OpenAI ammette : i modelli IA sono evasi dal sandbox e hanno attaccato Hugging Face hwupgrade.it · 6d 1h ago
99 OpenAI Says Its AI Models Escaped Sandbox, Targeted Hugging Face to Cheat Benchmark reddit.com · 5d 23h ago
98 Przestrzenna AI skróci dystans do klienta: pirxe wspiera odzieżowego giganta LPP mamstartup.pl · 5d 22h ago
96 Glow emerges from stealth at $1.2B valuation to challenge endpoint security in the AI era techcrunch.com · 6d ago
95 As Bitcoin breaches $66K its latest bottom signal trapped buyers in a 20% loss cryptoslate.com · 6d ago
83 A Classifier That Teaches Itself: Self-Improving, Frozen-gate Training (SIFT) for Dynamic Document Classification arxiv.org · 6d 1h ago
28 OpenAI’s latest AI agent escaped security controls and hacked a tech company washingtonpost.com · 5d 23h ago
28 A year in a foxhole: How one Ukrainian soldier survived beneath the front lines reuters.com · 6d ago
26 Show HN: Busymail – an email client that surfaces the mail that needs you busymail.app · 5d 22h ago
25 One Docker socket to rule them all: Escaping Codex, Cursor, and Gemini CLI pillar.security · 6d 1h ago
21 Bielik.ai: Community-built, open-source LLMs for Polish and European languages bielik.ai · 5d 23h ago
21 Flock Said Its Cameras Don’t Track People. Then a Reporter Proved Them Wrong gadgetreview.com · 5d 23h ago
20 Show HN: I mapped my AI coding setup – 90 of 103 installed skills never fire github.com · 5d 23h ago
19 Show HN: Forkbench – a native-Mac control room for running CLI coding agent forkbench.com · 5d 22h ago