CommentsStage 2 RR under review at EMSE. The accepted Stage 1 protocol is publicly archived on OSF (DOI: this https URL (https://doi.org/10.17605/OSF.IO/TCFJR) )
机构
*
Zhejiang University \& Northwestern Polytechnical University China
;
Northwestern Polytechnical University China
;
Nantong University China
;
Zhejiang University China
;
Monash University Australia
;
Singapore Management University Singapore
;
Zhejiang University \& Northwestern Polytechnical University
;
Northwestern Polytechnical University
;
Nantong University
;
Zhejiang University
;
Monash University
;
Singapore Management University
CommentsCode and data: this https URL (https://github.com/1549080929-debug/math_agent) Keywords: LLM verification; verification autonomy; completeness; ground truth; trustworthy AI Writing and implementation assisted by an AI language model; all experiments, data, and research decisions are the author's own
Commentsv2: substantially extended. Adds a second case study (rpart) reached through R's.Call() interface, and a five-phase prologue reconstructing the portion of R's C API the package uses so its original C compiles without R. v1 covered KernSmooth only. Title shortened
Commentsthis http URL (http://skillnet.openkg.cn/;) add SkillNet-Gym, a benchmark for evaluating skill retrieval, utilization, composition, and SkillNet-Fabric for task-specific skill routing through lightweight Wikis
Comments15 pages. Submitted to Formal Methods for Autonomous Systems 2026 (FMAS 2026). Reproducibility artefact: DOI https://doi.org/10.5281/zenodo.21981425