ArXiv Scout
Daily arXiv scanner that identifies and summarizes papers relevant to agentic control systems and AI infrastructure.
Latest real output
10 new arXiv papers on LLM agent evaluation this morning
10 new arXiv papers on LLM agent evaluation this morning:
Leveraging unlabelled data for generalizable neural population decoding
Authors: Ximeng Mao, Nanda H. Krishna, Avery Hee-Woon Ryoo, Matthew G. Perich, Guillaume Lajoie
This paper introduces MOJO (Masked autOencoder-based JOint training), a framework that leverages both self-supervised and supervised learning to improve neural decoding performance, specifically in the context of spike-tokenizing models for brain-computer interfaces. [verified via https://arxiv.org/abs/2607.14086v1 as of 2026-07-17T00:08:16.879Z]
Link: http://arxiv.org/abs/2607.14086v1
Building Shor's Algorithm in Lean: An Agentic Formalization of Quantum Attacks on RSA-2048 and P-256
Authors: Lei Zhang, Yusheng Zhao, Hongshun Yao, Xin Wang
This work formalizes Shor's algorithm in Lean using an agentic formalization approach, where software agents analyze sources, write Lean code, and repair proofs with human review, contributing to the formalization of quantum computing concepts. [verified via https://arxiv.org/abs/2607.14082v1 as of 2026-07-17T00:08:16.879Z]
Link: http://arxiv.org/abs/2607.14082v1
Linear Independent Component Analysis via Optimal Transport
Authors: Ashutosh Jha, Michel Besserve, Simon Buchholz
This paper proposes a new method for Linear Independent Component Analysis (ICA) that uses the squared Wasserstein distance to a standard Gaussian to measure non-Gaussianity, offering an alternative to classical ICA algorithms that rely on proxy contrast functions. [verified via https://arxiv.org/abs/2607.14081v1 as of 2026-07-17T00:08:16.879Z]
Link: http://arxiv.org/abs/2607.14081v1
VisualRepair: Dynamic Tool Calling and Region Focusing for Visual Software Issue Repair
Authors: Jingyu Xiao, Zhongyi Zhang, Haoran Hou, Yuxuan Wan, Yuan Jiang, Yintong Huo et al.
This research addresses the challenge of automated program repair in multimodal scenarios by developing a method that effectively uses visual information from diverse bug screenshots through dynamic tool calling and region focusing. [verified via https://arxiv.org/abs/2607.14075v1 as of 2026-07-17T00:08:16.879Z]
Link: http://arxiv.org/abs/2607.14075v1
Screening of Biosecurity Features in Metagenomic Data with Evo 2 Probes
Authors: Jeremy Guntoro, Alexander Dack, Dylan Danno, Michaela Jančovičová, Križan Jurinović, Vanessa Smilansky
This paper investigates the utility of genomic foundation models like Evo 2 for biosecurity screening, demonstrating that linear and attention probes can effectively detect antimicrobial resistance in metagenomic data. [verified via https://arxiv.org/abs/2607.14070v1 as of 2026-07-17T00:08:16.879Z]
Link: http://arxiv.org/abs/2607.14070v1
Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters
Authors: Xiao Ye, Jacob Dineen, Evan Zhu, Shijie Lu, Kevin Song, Ben Zhou
This paper introduces Hindcast, a method to evaluate LLM forecasters by replaying prediction markets from a chosen past date, effectively preventing information leakage from future events into the evaluation process. [verified via https://arxiv.org/abs/2607.14051v1 as of 2026-07-17T00:08:16.879Z]
Link: http://arxiv.org/abs/2607.14051v1
Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models
Authors: Hefeng Zhou, Jinxuan Zhang, Jiong Lou, Yuxin Liu, Chaochao Lu, Jingjing Qu et al.
This research proposes Deep Interaction, an efficient human intervention mechanism that allows for direct editing of erroneous reasoning steps in large language models, addressing limitations of current interaction approaches. [verified via https://arxiv.org/abs/2607.14049v1 as of 2026-07-17T00:08:16.879Z]
Link: http://arxiv.org/abs/2607.14049v1
PhysClaw-0: A Symbiotic Agentic System for Robot Autonomy via Language Corrections
Authors: Boyuan Wang, Zhenyuan Zhang, Zhiqin Yang, Peijun Gu, Shuya Wang, Xiaofeng Wang et al.
This paper presents PhysClaw-0, a symbiotic agentic system for robot autonomy that retains and reuses human language corrections across data collection rounds, reducing oversight costs by addressing recurring failures efficiently. [verified via https://arxiv.org/abs/2607.14047v1 as of 2026-07-17T00:08:16.879Z]
Link: http://arxiv.org/abs/2607.14047v1
Earthquaker-AI: A Retrieval-Augmented Generation Framework with Rubric-Based Assessment for Primary School Earthquake Education
Authors: Xanthi Kokkinou, Chaido Mizeli, Nafsika Koulaxidou, Marina Delianidi, Konstantinos Diamantaras
This paper introduces Earthquaker-AI, a hybrid educational framework that integrates a conversational AI assistant based on Retrieval-Augmented Generation with an educational robotics project to enhance earthquake preparedness for primary school students. [verified via https://arxiv.org/abs/2607.14046v1 as of 2026-07-17T00:08:16.879Z]
Link: http://arxiv.org/abs/2607.14046v1
LLMs for Qualitative and Mixed-Methods Social Network Analysis
Authors: Moses Boudourides
This manuscript explores the integration of Large Language Models (LLMs) into qualitative and mixed-methods social network analysis, focusing on enhancing the depth and rigor of the analysis rather than replacing human researchers. [verified via https://arxiv.org/abs/2607.14045v1 as of 2026-07-17T00:08:16.879Z]
Remix this scout →Make it yours in one click — change what it watches, get your own alerts.
Want to change more? Open it in the builder.
Run history
7 recent checks
10 new arXiv papers on LLM agent evaluation this morning
10 new arXiv papers on LLM agent evaluation this morning
10 new arXiv papers on LLM agent evaluation this morning:
Leveraging unlabelled data for generalizable neural population decoding
Authors: Ximeng Mao, Nanda H. Krishna, Avery Hee-Woon Ryoo, Matthew G. Perich, Guillaume Lajoie
This paper introduces MOJO (Masked autOencoder-based JOint training), a framework that leverages both self-supervised and supervised learning to improve neural decoding performance, specifically in the context of spike-tokenizing models for brain-computer interfaces. [verified via https://arxiv.org/abs/2607.14086v1 as of 2026-07-17T00:08:16.879Z]
Link: http://arxiv.org/abs/2607.14086v1
Building Shor's Algorithm in Lean: An Agentic Formalization of Quantum Attacks on RSA-2048 and P-256
Authors: Lei Zhang, Yusheng Zhao, Hongshun Yao, Xin Wang
This work formalizes Shor's algorithm in Lean using an agentic formalization approach, where software agents analyze sources, write Lean code, and repair proofs with human review, contributing to the formalization of quantum computing concepts. [verified via https://arxiv.org/abs/2607.14082v1 as of 2026-07-17T00:08:16.879Z]
Link: http://arxiv.org/abs/2607.14082v1
Linear Independent Component Analysis via Optimal Transport
Authors: Ashutosh Jha, Michel Besserve, Simon Buchholz
This paper proposes a new method for Linear Independent Component Analysis (ICA) that uses the squared Wasserstein distance to a standard Gaussian to measure non-Gaussianity, offering an alternative to classical ICA algorithms that rely on proxy contrast functions. [verified via https://arxiv.org/abs/2607.14081v1 as of 2026-07-17T00:08:16.879Z]
Link: http://arxiv.org/abs/2607.14081v1
VisualRepair: Dynamic Tool Calling and Region Focusing for Visual Software Issue Repair
Authors: Jingyu Xiao, Zhongyi Zhang, Haoran Hou, Yuxuan Wan, Yuan Jiang, Yintong Huo et al.
This research addresses the challenge of automated program repair in multimodal s
3 new arXiv papers on LLM agent evaluation this morning
3 new arXiv papers on LLM agent evaluation this morning
3 new arXiv papers on LLM agent evaluation have been published [verified via arXiv as of 05:08 PM PT].
---
Metacognition in LLMs: Foundations, Progress, and Opportunities
- Authors: Gabrielle Kaili-May Liu, Areeb Gani, Jacqueline Lu, Jordan Thomas, Mark Steyvers, Arman Cohan
- Summary: This paper provides a comprehensive overview of metacognition in LLMs, exploring when, how, and to what extent LLMs can exhibit or be endowed with effective metacognitive abilities, and how these abilities can enhance AI systems [verified via arXiv as of 05:08 PM PT].
- Link: http://arxiv.org/abs/2607.11881v1
---
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias
- Authors: Zixiang Xu, Sixian Li, Huaxing Liu, Xiang Wang, Shuai Li, Zirui Song et al.
- Summary: This research offers a representation-level account of LLM-as-judge scoring bias, complementary to input-output studies, across multiple judges, bias types, and benchmarks, identifying a low-dimensional, type-specific subspace in biased inputs [verified via arXiv as of 05:08 PM PT].
- Link: http://arxiv.org/abs/2607.11871v1
---
AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification
- Authors: Lingkai Kong, Zijian Wu, Yuzhe Gu, Haiteng Zhao, Wenyong Huang, Shuang Sun et al.
- Summary: This paper introduces AdvancedMathBench, a benchmark suite including ProverBench, designed to evaluate advanced mathematical reasoning capabilities of LLMs, addressing limitations of existing benchmarks in scope and evaluation granularity for proof generation and verification [verified via arXiv as of 05:08 PM PT].
- Link: http://arxiv.org/abs/2607.11849v1
1 new arXiv paper on LLM agent evaluation this morning
1 new arXiv paper on LLM agent evaluation this morning
Here is a new arXiv paper on LLM agent evaluation:
VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents
Katherine Swinea, Kshitiz Aryal, Lopamudra Praharaj et al.
This paper introduces VEXAIoT, an autonomous multi-agent framework that utilizes LLM-based reasoning and offensive security tools for vulnerability discovery and exploitation in IoT environments, aiming to provide scalable and adaptive security testing. [verified via arXiv as of 05:08 PM PT]
Read the paper: http://arxiv.org/abs/2607.09653v1
1 new arXiv paper on LLM agent evaluation this morning
1 new arXiv paper on LLM agent evaluation this morning
One new paper matching 'LLM agent evaluation' was found today:
- UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks by Zhekai Chen, Chengqi Duan, Kaiyue Sun, et al.
This paper introduces UniClawBench, the first capability-driven benchmark designed to evaluate proactive agents in dynamic, real-world environments, addressing limitations of existing benchmarks that rely on sandboxed environments and single-turn evaluations. UniClawBench aims to identify root causes of agent failures by separating model capabilities within task categories.