arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Science and Technology of China(中国科学技术大学)

共收录 2226
2603.09312 2026-03-11 cs.CV

IntroSVG: Learning from Rendering Feedback for Text-to-SVG Generation via an Introspective Generator-Critic Framework

IntroSVG: 通过反思生成器-批评者框架从渲染反馈中学习以实现文本到SVG生成

Feiyu Wang, Jiayuan Yang, Zhiyuan Zhao, Da Zhang, Bingyu Li, Peng Liu, Junyu Gao

机构 * Fudan University(复旦大学) TeleAI Northwestern Polytechnical University(西北工业大学) University of Science and Technology of China(中国科学技术大学)

AI总结 IntroSVG通过反思生成器-批评者框架,结合监督微调和直接偏好优化,实现文本到SVG生成的高质量输出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09259 2026-03-11 cs.CV cs.RO

Implicit Geometry Representations for Vision-and-Language Navigation from Web Videos

从网络视频中隐式几何表示的视觉-语言导航

Mingfei Han, Haihong Hao, Liang Ma, Kamila Zhumakhanova, Ekaterina Radionova, Jingyi Zhang, Xiaojun Chang, Xiaodan Liang, Ivan Laptev

机构 * Department of Computer Vision, Mohamed Bin Zayed University of Artificial Intelligence(计算机视觉系,Mohamed Bin Zayed人工智能大学) School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学) Shenzhen Campus of Sun Yat-Sen University(中山大学深圳校区)

AI总结 本研究提出基于网络视频的隐式几何表示方法,用于视觉-语言导航,通过提升数据利用率和鲁棒性,实现更广泛的应用。

Comments Extension of CVPR 2025 RoomTour3D with implicit geometric representations

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05871 2026-03-11 cs.CV

Pathwise Test-Time Correction for Autoregressive Long Video Generation

逐帧测试时校正用于自回归长视频生成

Xunzhi Xiang, Zixuan Duan, Guiyu Zhang, Haiyu Zhang, Zhe Gao, Junta Wu, Shaofeng Zhang, Tengfei Wang, Qi Fan, Chunchao Guo

机构 * Nanjing University(南京大学) Tencent Hunyuan(腾讯文言) Chinese University of Hong Kong Shenzhen(香港中文大学(深圳)) University of Science and Technology of China(中国科学技术大学)

AI总结 本文提出TTC方法,通过利用初始帧作为参考锚点,有效校正自回归长视频生成中的漂移问题,提升生成质量并延长生成长度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18811 2026-03-11 cs.CV cs.AI

Mitigating Long-Tail Bias in HOI Detection via Adaptive Diversity Cache

通过自适应多样性缓存缓解HOI检测中的长尾偏差

Yuqiu Jiang, Xiaozhen Qiao, Yifan Chen, Ye Zheng, Zhe Sun, Xuelong Li

机构 * College of Future Information Technology, Fudan University(复旦大学未来信息技术学院) School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学与技术学院) Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究院(TeleAI))

AI总结 本文提出自适应多样性缓存模块,通过无训练机制缓解HOI检测中的长尾偏差,提升稀有类别检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01290 2026-03-11 cs.LG cs.AI

Rating Quality of Diverse Time Series Data by Meta-learning from LLM Judgment

通过LLM判断对异质时间序列数据进行质量评级

Shunyu Wu, Dan Li, Wenjie Feng, Haozheng Ye, Jian Lou, See-Kiong Ng

机构 * Sun Yat-sen University(中山大学) State Key Laboratory of Al Safety(人工智能安全国家重点实验室) University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学)

AI总结 本文提出TSRating框架,利用LLM判断对异质时间序列数据进行质量评级,通过元学习提升跨领域适应性。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21147 2026-03-11 cs.LG

Semi-Supervised Conformal Prediction With Unlabeled Nonconformity Score

半监督置信预测与未标记非一致性评分

Xuanning Zhou, Zihao Shi, Hao Zeng, Xiaobo Xia, Bingyi Jing, Hongxin Wei

机构 * Southern University of Science and Technology(南方科技大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Fudan University(复旦大学) University of Science and Technology of China(中国科学技术大学) Shenzhen Loop Area Institute(深圳河套学院)

AI总结 本文提出SemiCP,通过结合标记和未标记数据提升置信预测的覆盖性能,实验表明在有限标记数据下可显著降低覆盖差距。

Comments Accept by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09084 2026-03-11 cs.CV

OmniEdit: A Training-free framework for Lip Synchronization and Audio-Visual Editing

OmniEdit: 一种无需训练的唇同步与音频视觉编辑框架

Lixiang Lin, Siyuan Jin, Jinshan Zhang

机构 * HiThink Research(HiThink研究机构) University of Science and Technology of China(中国科学技术大学) Zhejiang University(浙江大学)

AI总结 OmniEdit提出一种无需训练的框架,通过替换编辑序列和去除随机元素,实现唇同步与音频视觉编辑的高效稳定处理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08113 2026-03-10 cs.CV

SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving

SAMoE-VLA:一种面向自动驾驶的场景自适应混合专家视觉-语言-动作模型

Zihan You, Hongwei Liu, Chenxu Dang, Zhe Wang, Sining Ang, Aoqi Wang, Yan Wang

机构 * Institute for AI Industry Research (AIR), Tsinghua University(人工智能产业研究院(AIR),清华大学) School of Instrument Science and Engineering, Southeast University(仪器科学与工程学院,东南大学) Zhili College, Tsinghua University(紫荆学院,清华大学) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学) Department of Automation, University of Science and Technology of China(自动化学院,中国科学技术大学) Department of Automation, University of Science and Technology Beijing(自动化学院,北京科技大学)

AI总结 SAMoE-VLA通过场景自适应混合专家机制提升自动驾驶中的视觉-语言-动作推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08035 2026-03-10 cs.AI cs.LG

CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling

CDRRM:基于对比的评分标准生成用于可靠和可解释的奖励建模

Dengcan Liu, Fengkai Yang, Xiaohan Wang, Shurui Yan, Jiajun Chai, Jiahao Li, Yikun Ban, Zhendong Mao, Wei Lin, Guojun Yin

机构 * University of Science and Technology of China(科学技术大学) Peking University(北京大学) BeiHang University(北航大学)

AI总结 CDRRM通过对比驱动的评分标准生成方法,实现可靠且可解释的奖励建模,有效缓解评估偏见并提升数据效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08034 2026-03-10 cs.CV cs.AI

Solution to the 10th ABAW Expression Recognition Challenge: A Robust Multimodal Framework with Safe Cross-Attention and Modality Dropout

解决第10届ABAW表情识别挑战的方案:一种具有安全交叉注意力和模态dropout的鲁棒多模态框架

Jun Yu, Naixiang Zheng, Guoyuan Wang, Yunxiang Zhang, Lingsi Zhu, Jiaen Liang, Wei Huang, Shengping Liu

机构 * University of Science and Technology of China(中国科学技术大学) Unisound AI Technology Co., Ltd.(Unisound人工智能技术有限公司)

AI总结 本文提出一种鲁棒多模态框架,通过安全交叉注意力和模态dropout处理现实环境中的遮挡和缺失模态问题,提升情绪识别性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07897 2026-03-10 cs.LG

LeJOT-AutoML: LLM-Driven Feature Engineering for Job Execution Time Prediction in Databricks Cost Optimization

LeJOT-AutoML:基于LLM的特征工程用于Databricks成本优化中的作业执行时间预测

Lizhi Ma, Yi-Xiang Hu, Yihui Ren, Feng Wu, Xiang-Yang Li

机构 * University of Science and Technology of China(中国科学技术大学) Lenovo(联想)

AI总结 LeJOT-AutoML利用LLM驱动的AutoML框架,通过动态生成特征提升Databricks作业执行时间预测的准确性,从而实现成本优化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07671 2026-03-10 cs.LG

Beyond Surrogates: A Quantitative Analysis for Inter-Metric Relationships

超越替代方案:跨度量关系的定量分析

Yuanhao Pu, Defu Lian, Enhong Chen

机构 * School of Artificial Intelligence & Data Science, University of Science and Technology of China(人工智能与数据科学学院,中国科学技术大学) School of Computer Science & Technology, University of Science and Technology of China(计算机科学与技术学院,中国科学技术大学) State Key Laboratory of Cognitive Intelligence, China(认知智能国家重点实验室,中国)

AI总结 本文提出统一理论框架,定量分析指标间关系,解决指标不匹配问题,确保离线改进与在线目标一致。

Comments 18 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07365 2026-03-10 cs.SD cs.AI cs.CL cs.MM eess.AS

Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning

多领域音频问答基准:面向声音内容推理

Chao-Han Huck Yang, Sreyan Ghosh, Qing Wang, Jaeyeon Kim, Hengyi Hong, Sonal Kumar, Guirui Zhong, Zhifeng Kong, S Sakshi, Vaibhavi Lokegaonkar, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha, Gunhee Kim, Jun Du, Rafael Valle, Bryan Catanzaro

机构 * NVIDIA University of Maryland, College Park(马里兰大学 College Park 分校) University of Science and Technology of China(中国科学技术大学) Seoul National University(首尔国立大学) Adobe

AI总结 DCASE 2025挑战赛提出多领域音频问答基准,通过生物声学、时间声音景观和复杂问答子集测试音频-语言模型在多样声音场景中的交互式问答能力,推动音频理解和推理能力发展。

Comments Dataset: https://huggingface.co/datasets/PeacefulData/2025_DCASE_AudioQA_Official DCASE Task-5 challenge: dcase.community/challenge2025/task-audio-question-answering. Accepted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11093 2026-03-10 quant-ph cs.LG

Simulating Non-Markovian Open Quantum Dynamics with Neural Quantum States

用神经量子态模拟非马尔可夫开放量子动力学

Long Cao, Liwei Ge, Daochi Zhang, Xiang Li, Yao Wang, Rui-Xue Xu, YiJing Yan, Xiao Zheng

机构 * Hefei National Research Center for Physical Sciences at the Microscale, University of Science and Technology of China(合肥微尺度物质科学国家研究中心,中国科学技术大学) Department of Chemistry, Fudan University(复旦大学化学系) Hefei National Laboratory(合肥国家实验室)

AI总结 本文提出一种基于神经量子态和耗散子嵌入量子主方程的方法,用于高效模拟非马尔可夫开放量子动力学,提升计算可扩展性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07107 2026-03-10 cs.IR cs.AI

Efficient Personalized Reranking with Semi-Autoregressive Generation and Online Knowledge Distillation

高效个性化重排序:半自动生成与在线知识蒸馏

Kai Cheng, Hao Wang, Wei Guo, Weiwen Liu, Yong Liu, Yawen Li, Enhong Chen

机构 * University of Science and Technology of China(中国科学技术大学) Huawei(华为) Shanghai Jiao Tong University(上海交通大学) Beijing University of Posts and Telecommunications(北京邮电大学)

AI总结 本文提出PSAD框架,通过半自动生成与在线知识蒸馏提升个性化重排序的效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07078 2026-03-10 cs.AI cs.CL

CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs

CoTJudger: 一种基于图的框架,用于自动评估链式推理效率和冗余性在LRMs中

Siyi Li, Jiajun Shi, Shiwen Ni, Ge Zhang, Shuaimin Li, Shijian Wang, Zhoufutu Wen, Yizhi Li, Hamid Alinejad-Rokny, Jiaheng Liu, Min Yang, Wenhao Huang

机构 * University of Science and Technology of China(科学技术大学) Shenzhen University of Advanced Technology(深圳先进技术大学) Shenzhen Institutes of Advanced Technology, CAS(深圳先进技术研究所,中国科学院) Southeast University(东南大学) Nanjing University(南京大学) Beihang University(北航) University of Manchester(曼彻斯特大学)

AI总结 CoTJudger通过构建依赖图提取最短有效路径,评估链式推理的效率与冗余,揭示模型中的冗余问题及失败模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13695 2026-03-10 cs.AI math.AC math.CO math.CT

Can a Lightweight Automated AI Pipeline Solve Research-Level Mathematical Problems?

能否一个轻量级的自动化AI流水线解决研究级别的数学问题?

Lve Meng, Weilong Zhao, Yanzhi Zhang, Haoxiang Guan, Jiyan He

机构 * University of Science and Technology of China(中国科学技术大学) Université Paris Cité(巴黎cité大学) Zhongguancun Academy(中关村学院)

AI总结 本研究展示了一种轻量级自动化AI流水线,能够解决复杂的研究级数学问题,并通过验证和开源实现。

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09486 2026-03-10 cs.CV cs.AI cs.MM

Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding

视频-EM:面向长视频理解的事件中心型片段记忆

Yun Wang, Long Zhang, Jingren Liu, Jiaqi Yan, Zhanjie Zhang, Jiahao Zheng, Ao Ma, Run Ling, Xun Yang, Dapeng Wu, Xiangyu Chen, Xuelong Li

机构 * City University of Hong Kong(香港城市大学) University of Science and Technology of China(中国科学技术大学) Tianjin University(天津大学) Nanjing University(南京大学) Zhejiang University(浙江大学) The Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究所(TeleAI))

AI总结 Video-EM通过事件中心型片段记忆框架,将长视频问答转化为事件构建与记忆细化,提升长视频理解的连贯性与可靠性。

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13347 2026-03-10 cs.CV

$π^3$: Permutation-Equivariant Visual Geometry Learning

π³:排列等变视觉几何学习

Yifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang, Yang Zhou, Zizun Li, Junyi Chen, Jiangmiao Pang, Chunhua Shen, Tong He

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院) Zhejiang University(浙江大学) University of Science and Technology of China(中国科学技术大学) Fudan University(复旦大学)

AI总结 π³通过排列等变架构实现无需参考视角的视觉几何重建,提升相机姿态和点地图重建的准确性和鲁棒性。

Comments Project page: https://yyfz.github.io/pi3/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06213 2026-03-09 cs.CV cs.AI

Cut to the Chase: Training-free Multimodal Summarization via Chain-of-Events

直击核心:一种无需训练的多模态摘要方法 via 事件链

Xiaoxing You, Qiang Huang, Lingyu Li, Xiaojun Chang, Jun Yu

机构 * School of Computer Science, Hangzhou Dianzi University(杭州电子科技大学计算机科学学院) School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)智能科学与工程学院) School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学与技术学院)

AI总结 无需训练的多模态摘要方法CoE通过事件链和层次事件图实现跨模态整合与时间推理,提升摘要质量与领域适应性。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00543 2026-03-09 cs.CV

Cross-Scale Pansharpening via ScaleFormer and the PanScale Benchmark

跨尺度 pansharpening 通过 ScaleFormer 和 PanScale 数据集

Ke Cao, Xuanhua He, Xueheng Li, Lingting Zhu, Yingying Wang, Ao Ma, Zhanjie Zhang, Man Zhou, Chengjun Xie, Jie Zhang

机构 * HFIPS, Chinese Academy of Sciences(中国科学院HFIPS) University of Science and Technology of China(中国科学技术大学) The Hong Kong University of Science and Technology(香港科学与技术大学) The University of Hong Kong(香港大学) Xiamen University(厦门大学) Zhejiang University(浙江大学)

AI总结 本文提出 ScaleFormer 架构和 PanScale 数据集,通过跨尺度 pansharpening 方法提升多尺度场景下的图像融合质量与泛化能力。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07945 2026-03-09 cs.LG

One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning

一个模型用于所有任务:利用高效的world models进行多任务规划

Yuan Pu, Yazhe Niu, Jia Tang, Junyu Xiong, Shuai Hu, Hongsheng Li

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong MMLab(香港中文大学 MMLab) University of Science and Technology of China(中国科学技术大学) Novosibirsk State University(新西伯利亚州立大学) Centre for Perceptual and Interactive Intelligence(感知与交互智能中心)

AI总结 ScaleZero通过混合专家架构和动态参数扩展策略,实现单一模型在多任务规划中的高效性能。

Comments 55 pages, 20 figures. Accepted as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05240 2026-03-06 cs.AI

GCAgent: Enhancing Group Chat Communication through Dialogue Agents System

GCAgent:通过对话代理系统增强群聊交流

Zijie Meng, Zheyong Xie, Zheyu Ye, Chonggang Lu, Zuozhu Liu, Zihan Niu, Yao Hu, Shaosheng Cao

机构 * Zhejiang University(浙江大学) Xiaohongshu Inc.(小红书公司) University of Science and Technology of China(中国科学技术大学)

AI总结 GCAgent通过集成娱乐和实用导向的对话代理系统,提升群聊交流效果,实现多参与者对话的高效管理与增强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03137 2026-03-06 cs.RO

RL-Based Coverage Path Planning for Deformable Objects on 3D Surfaces

基于强化学习的可变形物体3D表面覆盖路径规划

Yuhang Zhang, Jinming Ma, Feng Wu

机构 * School of Computer Science and Technology, University of Science and Technology of China(计算机科学与技术学院,中国科学技术大学) Xiaomi Robotics Lab(小米机器人实验室)

AI总结 本文提出基于强化学习的覆盖路径规划方法,利用谐波UV映射和SGCNN处理可变形物体的触觉反馈,通过模拟器训练机器人高效完成表面擦拭任务。

Comments 8 pages, 8 figures. Accepted to the 2026 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03604 2026-03-06 cs.AI

Interleaved Tool-Call Reasoning for Protein Function Understanding

交错工具调用推理用于蛋白质功能理解

Chuanliu Fan, Zicheng Ma, Huanran Meng, Aijia Zhang, Wenjie Du, Jun Zhang, Yi Qin Gao, Ziqiang Cao, Guohong Fu

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Institute of Artificial Intelligence, Soochow University(苏州大学人工智能研究院) Changping Laboratory(昌平实验室) School of Software Engineering, USTC(中国科学技术大学软件学院)

AI总结 PFUA通过整合领域特定工具和可验证中间证据,提高了蛋白质功能预测的性能,平均提升达103%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22796 2026-03-06 cs.CV

Parallel Diffusion Solver via Residual Dirichlet Policy Optimization

并行扩散求解器 via 剩余狄利克雷策略优化

Ruoyu Wang, Ziyu Li, Beier Zhu, Liangyu Yuan, Hanwang Zhang, Xun Yang, Xiaojun Chang, Chi Zhang

机构 * AGI lab, Westlake University(西溪大学AGI实验室) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Nanyang Technological University(南洋理工大学) Shanghai Jiao Tong University(上海交通大学) University of Science and Technology of China(中国科学技术大学)

AI总结 本文提出EPD-Solver,通过并行梯度评估和狄利克雷策略优化,提升扩散模型的低延迟采样性能。

Comments arXiv admin note: substantial text overlap with arXiv:2507.14797

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06690 2026-03-06 cs.CL

Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form Generation

在生成过程中思考:面向个性化长文本生成的即兴推理

Chengbing Wang, Yang Zhang, Wenjie Wang, Xiaoyan Zhao, Fuli Feng, Xiangnan He, Tat-Seng Chua

机构 * University of Science and Technology of China(科学技术大学) National University of Singapore(国立新加坡大学) The Chinese University of Hong Kong(香港中文大学) National Engineering Laboratory for BITA, University of Science and Technology of China(BITA国家工程实验室,科学技术大学)

AI总结 FlyThinker通过在生成过程中进行动态推理,提升个性化长文本生成的效率和效果。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18864 2026-03-06 cs.CV

Flatness Guided Test-Time Adaptation for Vision-Language Models

基于平坦度引导的视觉-语言模型测试时适应

Aodi Li, Liansheng Zhuang, Xiao Long, Houqiang Li, Shafei Wang

机构 * School of Cyber Science and Technology, University of Science and Technology of China(中国科学技术大学网络科学与技术学院) National Engineering Laboratory for Brain-Inspired Intelligence Technology and Applications, University of Science and Technology of China(中国科学技术大学脑启发式智能技术与应用国家工程实验室) Peng Cheng Laboratory, Shenzhen, China(深圳鹏城实验室)

AI总结 本文提出基于平坦度引导的视觉-语言模型测试时适应框架,通过训练时的平坦最小值与测试时损失景观的对齐,提升模型在分布偏移下的适应性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03985 2026-03-05 cs.CV

RIVER: A Real-Time Interaction Benchmark for Video LLMs

RIVER:面向视频大语言模型的实时交互基准

Yansong Shi, Qingsong Zhao, Tianxiang Jiang, Xiangyu Zeng, Yi Wang, Limin Wang

机构 * School of Information Science And Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) College of Computer Science and Artificial Intelligence, Fudan University(计算机科学与人工智能学院,复旦大学) State Key Lab of Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学)

AI总结 RIVER基准通过引入实时交互任务框架,改进视频大语言模型的实时交互能力,提升长期记忆与未来感知表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03960 2026-03-05 cs.RO cs.CV

Structural Action Transformer for 3D Dexterous Manipulation

结构动作变换器用于3D灵巧操作

Xiaohan Lei, Min Wang, Bohong Weng, Wengang Zhou, Houqiang Li

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知国家重点实验室,中国科学技术大学) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(人工智能研究院,合肥综合性国家科学中心)

AI总结 本文提出结构动作变换器,通过结构化视角解决高自由度机械手的跨身体技能转移问题,实现更高效的灵巧操作。

Comments Accepted by CVPR

详情

展开后加载摘要…

URL PDF HTML 收藏