Octopi: Object Property Reasoning with Large Tactile-Language Models
专题命中 推理评测 :reasoning(title,abstract)
Comments Accepted at Robotics: Science and Systems (R:SS 2024)
AI 大模型
大模型数学、逻辑、规划、多步推理和测试时计算能力。
专题命中 推理评测 :reasoning(title,abstract)
Comments Accepted at Robotics: Science and Systems (R:SS 2024)
专题命中 推理评测 :reasoning(title,abstract)
Comments Accepted by ICML 2024
专题命中 推理评测 :reasoning(title,abstract)
Comments Under review
专题命中 推理评测 :reasoning(title,abstract)
Comments Technical report
专题命中 推理评测 :reasoning(title,abstract)
Comments Code, models, and data are available at https://github.com/dvlab-research/LISA
专题命中 推理评测 :reasoning(title,abstract)
Comments To be appeared at 2024 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER)
专题命中 推理评测 :reasoning(title,abstract)
Comments 5 pages, 2 figures , in proceeding of 5th International Seminar on Artificial Intelligence, Networking and Information Technology
专题命中 推理评测 :reasoning(title,abstract)
专题命中 推理评测 :reasoning(title,abstract)
Comments 8 pages, 4 figures
专题命中 推理评测 :reasoning(title,abstract)
Comments Accepted by the 38th Annual AAAI Conference on Artificial Intelligence (AAAI 2024) in December 2023
专题命中 推理评测 :reasoning(title,abstract)
专题命中 推理评测 :chain-of-thought(title,abstract)
Comments 10 pages (not including references and appendix), 14 figures (7 in main paper, 7 in appendix); (v3) camera-ready version
专题命中 推理评测 :reasoning(title,abstract)
专题命中 推理评测 :reasoning(title,abstract)
专题命中 推理评测 :CoT(title,abstract)
专题命中 推理评测 :reasoning(title,abstract)
专题命中 推理评测 :reasoning(title,abstract)
专题命中 推理评测 :reasoning(title,abstract)
专题命中 推理评测 :reasoning(title,abstract)
Comments The paper is under consideration at Computer Vision and Image Understanding
专题命中 推理评测 :reasoning(title,abstract)
专题命中 推理评测 :reasoning(title,abstract)
Comments CVPR2022; Code, data: https://github.com/leonnnop/VAR
专题命中 推理评测 :reasoning(title,abstract)
Comments 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2021)
专题命中 推理评测 :reasoning(title,abstract)
专题命中 推理评测 :reasoning(title,abstract)
Comments Presented at ICPR 2020
人工智能代理知道任务何时简单吗?迈向复杂性感知推理与执行
专题命中 推理评测 :reasoning(title,comments);分类 cs.CL、cs.AI
AI总结 研究探讨LLM代理缺乏任务感知执行范围估计能力,提出E3方法,在MSE-Bench基准测试及真实模型工具中验证,该方法能在保证成功率时大幅降低成本,推动迈向工程基础人工智能,还发布了框架和基准。
Comments 27 pages, 8 figures, 8 tables. Code and benchmark: https://github.com/eejyin/Do-AI-Agents-Know-When-a-Task-Is-Simple-Toward-Complexity-Aware-Reasoning-and-Execution
LAUDE: 基于LLM的硬件设计单元测试生成与调试
机构 * Dept. of Electrical and Computer Engineering, University of Illinois Chicago(伊利诺伊大学芝加哥分校电子与计算机工程系) ; Microsoft(微软公司)
专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.AI
AI总结 LAUDE利用大语言模型能力,实现硬件设计的单元测试生成与调试,有效检测和修复设计中的错误。
Comments 11 Pages, 9 Figures, 9 Tables Submitted to AAAI 2027
Gaokerena:一个小型波斯语医疗语言模型家族
机构 * University of Isfahan(伊斯法罕大学) ; University of Tehran(德黑兰大学) ; University of Windsor(温莎大学) ; Alzahra University(阿勒扎哈拉大学) ; University of Texas at Dallas(德克萨斯大学达拉斯分校)
专题命中 推理评测 :chain-of-thought(abstract,abstract_cn);reasoning(abstract);分类 cs.CL
AI总结 针对波斯语医疗语言模型研究不足的问题,该研究推出Gaokerena家族,包含Gaokerena-V和Gaokerena-R两款模型,在医疗问答基准上取得性能提升,还配备不确定性头,为本地化数字医疗提供基础但仍需完善。
Comments 29 pages, 9 figures
比较还是不比较:论评估社会偏见的方法论实践
机构 * Tsinghua University(清华大学) ; The Hebrew University of Jerusalem(耶路撒冷希伯来大学) ; Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, TU Darmstadt and National Research Center for Applied Cybersecurity ATHENE(普适知识处理实验室(UKP Lab),计算机科学系,达姆施塔特工业大学及国家应用网络安全研究中心ATHENE)
专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL
AI总结 提出统一框架标准化异构基准,揭示孤立评估与强制比较设置间的系统性范式差距,发现比较设置会催化潜在歧视,且链式思维推理加剧偏见,规模越大偏见越强。
ThinkDeception: 一种用于可解释多模态欺骗检测的渐进式强化学习框架
机构 * Xi'an Jiaotong-Liverpool University(西安交通大学利物浦大学)
专题命中 推理评测 :CoT(abstract,abstract_cn);reasoning(abstract);分类 cs.AI
AI总结 提出ThinkDeception框架,将多模态大语言模型引入欺骗检测,通过逐步推理和视觉-音频一致性组相对策略优化(VAC-GRPO)实现可解释的认知推理,在主流基准上达到新SOTA。
Comments 10pages,4figures
用大型语言模型模拟学生的Java编程错误
机构 * University of California, Los Angeles(加州大学洛杉矶分校)
专题命中 推理评测 :CoT(abstract,abstract_cn);chain-of-thought(abstract);分类 cs.CL
AI总结 探索用LLM模拟学生编程错误,评估五种模型在多样性和对齐性上的表现,发现Claude Sonnet 4平衡最佳,且合成错误与真实错误难以区分。