100 Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization arxiv.org · 1h 6m ago
100 Humanly: A Configurable and Traceable Environment for Human-AI Collaborative Writing arxiv.org · 1h 6m ago
100 Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models arxiv.org · 1h 6m ago
100 Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings arxiv.org · 1h 6m ago
100 Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study arxiv.org · 1h 6m ago
100 Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMs arxiv.org · 1h 6m ago
100 Developing and Validating the Spanish Version of the Large Language Models Dependency Scale (LLM-D12-SP) arxiv.org · 1h 6m ago
100 From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models arxiv.org · 1h 6m ago
100 Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity arxiv.org · 1h 6m ago
100 Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills arxiv.org · 1h 6m ago
100 From Obligation to Specification: A Survey on Validating EU AI Act Requirements in RE arxiv.org · 1h 6m ago
100 Codec-Gauge: Learning Compression-Friendly Gauges for Transformer KV Caches arxiv.org · 3d 1h ago
100 From Atoms to Entropy: Optimal Noise Allocation for Diffusion Training in the Convex Regime arxiv.org · 3d 1h ago
100 HypNO: A Graph-Based Neural Operator with Physics-Informed Message Passing for Hyperbolic Conservation Laws arxiv.org · 3d 1h ago
100 Conflict Resolution under Degraded Surveillance in Air Corridors Using Multi-Agent Reinforcement Learning arxiv.org · 3d 1h ago
100 SevDiff: Severity-Conditioned Diffusion for Long-Tail Conflict Trajectory Generation arxiv.org · 3d 1h ago
100 Beyond SBDD: Geometric Deep Learning in Polypharmacology and Multi-target Drug Design arxiv.org · 3d 1h ago
100 SenCos-GEM: SENet-Calibrated and Law-of-Cosines-Constrained Geometry-Enhanced Molecular Representation for Property Prediction arxiv.org · 3d 1h ago
100 ClickGuard: Detecting and Spoiling Clickbait News with Informativeness Measures and Large Language Models arxiv.org · 3d 1h ago
100 Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs arxiv.org · 3d 1h ago
100 InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents arxiv.org · 3d 1h ago
100 PlanE: Meta Planning of Data, Tuning, and Inference for Extractive-based LLMs arxiv.org · 3d 1h ago
100 VeriSimpl: Robust Optimization Modeling from Natural Language using Simplification-based Verification arxiv.org · 3d 1h ago
100 Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment arxiv.org · 3d 1h ago
100 Beyond Liars’ Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs arxiv.org · 3d 1h ago
100 Enabling Scalable Topology Inference in Distribution Systems via Constrained Multi-Source Inference arxiv.org · 3d 1h ago
100 Routing Without Training: Controllable-Ratio LLM Offloading via Reliability Gating arxiv.org · 3d 1h ago
100 The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path arxiv.org · 3d 1h ago
100 OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining arxiv.org · 3d 1h ago
100 Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering arxiv.org · 3d 1h ago
100 Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with Capability, State, and Shadow Invariants arxiv.org · 3d 1h ago
100 DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making arxiv.org · 3d 1h ago
100 AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics arxiv.org · 3d 1h ago
79 Codex with GPT 5.6 Sol Ultra is a powerhouse, and doing things i never thought possible this early. reddit.com · 3d 3h ago
36 AI will not trigger employment collapse, staffing company Adecco Group says reuters.com · 3d 2h ago
31 Execution-Free Agentic Program Repair for Enterprise-Scale Development [pdf] dl.acm.org · 3d 1h ago
31 Uncle Bob: My current strategy is to not read any code written by my agents xcancel.com · 3d 1h ago