arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Science and Technology of China(中国科学技术大学)

共收录 2226
2604.20366 2026-04-23 cs.CV

Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation

在不降低性能的情况下缓解大型视觉-语言模型的幻觉

Xingyu Zhu, Junfeng Fang, Shuo Wang, Beier Zhu, Zhicai Wang, Yonghui Yang, Xiangnan He

机构 * MoE Key Lab of BIPC(BIPC摩尔实验室) University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学)

AI总结 本文提出MPD框架,通过语义感知组件解耦和可解释参数更新,有效减少幻觉并保持生成能力,实验显示在LLaVA-Bench和MME上减少23.4%幻觉同时保持97.4%的生成能力。

Comments ACL 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20261 2026-04-23 cs.AI

Memory-Augmented LLM-based Multi-Agent System for Automated Feature Generation on Tabular Data

基于记忆增强的LLM多智能体系统用于表格数据自动特征生成

Fengxian Dong, Zhi Zheng, Xiao Han, Wei Chen, Jingqing Ruan, Tong Xu, Yong Chen, Enhong Chen

机构 * University of Science and Technology of China(中国科学技术大学) Zhejiang University of Technology(浙江工业大学) Meituan(美团)

AI总结 本文提出MALMAS系统,通过多智能体与记忆模块提升表格数据自动特征生成的效率与质量,实验验证其有效性。

Comments 16 pages (including appendix), 4 main figures, 15 tables. Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19544 2026-04-22 cs.AI

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling

DT2IT-MRM:去偏偏好构建与迭代训练用于多模态奖励建模

Zhihong Zhang, Jie Zhao, Xiaojian Huang, Jin Xu, Zhuodong Luo, Xin Liu, Jiansheng Wei, Xuejin Chen

机构 * University of Science and Technology of China(中国科学技术大学)

AI总结 本文提出DT2IT-MRM方法,通过去偏偏好构建、文本到图像偏好数据重构和迭代训练框架,解决多模态奖励建模中偏好数据的偏差、噪声和不一致问题,并在三个基准测试中取得新突破。

Comments code will be uploaded to https://github.com/zhang123434/DT2IT-MRM

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05301 2026-04-22 cs.CV

SmokeGS-R: Physics-Guided Pseudo-Clean 3DGS for Real-World Multi-View Smoke Restoration

SmokeGS-R:基于物理的伪清洁3DGS用于真实世界多视角烟雾修复

Xueming Fu, Lixia Han

机构 * University of Science and Technology of China(中国科学技术大学) Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

AI总结 本文提出SmokeGS-R,通过分离几何恢复与外观校正,利用改进的暗通道先验和引导滤波生成物理引导的伪清洁监督,训练纯清洁3D高斯点散布源模型,并通过几何均值参考聚合、LAB空间Reinhard传输和光高斯平滑和谐化其渲染结果,实现真实世界多视角烟雾修复。

Comments Lab Report for NTIRE 2026 3DRR Track 2

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03540 2026-04-22 cs.RO

Drift-Based Policy Optimization: Native One-Step Policy Learning for Online Robot Control

基于漂移的策略优化:面向在线机器人控制的原生一步生成策略

Yuxuan Gao, Yedong Shen, Shiqi Zhang, Wenhao Yu, Yifan Duan, Jia pan, Jiajia Wu, Jiajun Deng, Yanyong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) iFLYTEK

AI总结 本文提出一种原生一步生成策略框架,通过将迭代细化移至训练阶段,提升在线机器人控制的效率与稳定性,实验显示其在推理速度和性能上优于多步扩散策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18871 2026-04-22 cs.CV cs.AI cs.CL

OmniGen2: Towards Instruction-Aligned Multimodal Generation

OmniGen2:面向指令对齐的多模态生成

Chenyuan Wu, Pengfei Zheng, Ruiran Yan, Shitao Xiao, Xin Luo, Yueze Wang, Wanli Li, Xiyan Jiang, Yexin Liu, Junjie Zhou, Ze Liu, Ziyi Xia, Chaofan Li, Haoge Deng, Jiahao Wang, Kun Luo, Bo Zhang, Defu Lian, Xinlong Wang, Zhongyuan Wang, Tiejun Huang, Zheng Liu

机构 * Beijing Academy of Artificial Intelligence(北京人工智能研究院) University of Science and Technology of China(中国科学技术大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Zhejiang University(浙江大学)

AI总结 OmniGen2通过双解码路径和解耦图像分词器,实现文本和图像生成的统一解决方案,提升多模态生成任务的性能与一致性,同时提供开源模型和数据集支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19083 2026-04-22 cs.CR cs.AI

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety

ProjLens: 揭示项目器在多模态模型安全中的作用

Kun Wang, Cheng Qian, Miao Yu, Lilan Peng, Liang Lin, Jiaming Zhang, Tianyu Zhang, Yu Cheng, Yang Wang

机构 * University of Science and Technology of China(中国科学技术大学) Beijing University of Aeronautics and Astronautics(北京航空航天大学) Nanyang Technological University(南洋理工大学) Southwest Jiaotong University(西南交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

AI总结 ProjLens通过分析多模态大语言模型中的后门攻击机制,揭示了项目器在安全漏洞中的关键作用,发现后门注入参数编码于低秩子空间,并通过实验验证了激活机制的差异。

Comments 18 pages ,15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18988 2026-04-22 cs.CV

A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation

一种具有结构化推理和反思细化的多智能体框架用于多模态共情响应生成

Liping Wang, Cheng Ye, Weidong Chen, Peipei Song, Bo Hu, Zhendong Mao

机构 * School of Information Science and Technology, USTC(信息科学与技术学院,中国科学技术大学)

AI总结 本文提出多智能体框架,通过结构化推理和反思细化提升多模态共情响应生成的准确性与情感感知能力,实验表明优于现有方法。

Comments Submitted to ACM Multimetida 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07761 2026-04-22 cs.AI cs.LG

TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards

TROJail: 多轮对话中针对大语言模型的多轮攻击优化

Xiqiao Xiong, Ouxiang Li, Zhuo Liu, Moxin Li, Wentao Shi, Fengbin Zhu, Qifan Wang, Fuli Feng

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) Meta AI

AI总结 本文提出TROJail,通过多轮强化学习优化攻击策略,引入过程奖励提升攻击成功率,实验显示在多个模型和基准上效果显著。

Comments Accepted to ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18459 2026-04-21 cs.CV cs.AI

Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions

渐进式在线视频理解与证据对齐的透明决策

Kecheng Zhang, Zongxin Yang, Mingfei Han, Haihong Hao, Yunzhi Zhuge, Changlin Li, Junhan Zhao, Zhihui Li, Xiaojun Chang

机构 * University of Science and Technology of China(中国科学技术大学) Harvard University(哈佛大学) Department of CV, MBZUAI(MBZUAI计算机视觉部门) Dalian University of Technology(大连理工大学) Stanford University(斯坦福大学) Chicago University(芝加哥大学)

AI总结 本文提出一种新型框架,通过分离推理控制与记忆整合,解决在线视频理解中的决策透明、时间对齐和计算预算限制问题,显著提升了性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18348 2026-04-21 cs.CV cs.AI

AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generation

AdaCluster:用于视频生成中稀疏注意力的自适应查询-键聚类

Haoyue Tan, Shengnan Wang, Yulin Qiao, Juncheng Zhang, Youhui Bai, Ping Gong, Zewen Jin, Cheng Li

机构 * University of Science and Technology of China(中国科学技术大学) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院) Independent Researcher(独立研究员) University of Macau(澳门大学) The Chinese University of Hong Kong(香港中文大学)

AI总结 针对视频扩散变换器的高推理延迟问题,提出AdaCluster框架,通过自适应聚类方法提升生成速度并保持精度。

Comments CVPR 2026 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18240 2026-04-21 cs.AI

AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation

AJ-Bench:用于环境感知评估的代理作为裁判基准测试

Wentao Shi, Yu Wang, Yuyang Zhao, Yuxin Chen, Fuli Feng, Xueyuan Hao, Xi Su, Qi Gu, Hui Su, Xunliang Cai, Xiangnan He

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) Meituan(美团)

AI总结 AJ-Bench通过三个领域155个任务评估代理作为裁判的能力,揭示了基于代理的验证方法在信息获取和过程验证中的性能提升及挑战。

Comments Accepted to ACL 2026 Findings. 43 pages total, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18176 2026-04-21 cs.AI quant-ph

QuantumQA: Enhancing Scientific Reasoning via Physics-Consistent Dataset and Verification-Aware Reinforcement Learning

量子QA:通过物理一致的数据集和验证意识强化学习增强科学推理

Songxin Qu, Tai-Ping Sun, Yun-Jie Wang, Huan-Yu Liu, Cheng Xue, Xiao-Fan Xu, Han Fang, Yang Yang, Yu-Chun Wu, Guo-Ping Guo, Zhao-Yun Chen

机构 * Institute of Advanced Technology, University of Science and Technology of China(中国科学技术大学先进技术研究院) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院) School of Physics, University of Science and Technology of China(中国科学技术大学物理学院) School of Electronics and Information Engineering, Anhui University(安徽大学电子与信息工程学院) School of Computing, National University of Singapore(新加坡国立大学计算机学院)

AI总结 本文提出量子QA数据集和验证意识奖励模型,通过物理一致的数据集和强化学习提升科学推理能力,实验证明其在科学领域表现优异。

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18047 2026-04-21 cs.CV

GS-STVSR: Ultra-Efficient Continuous Spatio-Temporal Video Super-Resolution via 2D Gaussian Splatting

GS-STVSR:通过二维高斯散射实现超高效的连续时空视频超分辨率

Mingyu Shi, Xin Di, Long Peng, Boxiang Cao, Anran Wu, Zhanfeng Feng, Jiaming Guo, Renjing Pei, Xueyang Fu, Yang Cao, Zhengjun Zha

机构 * University of Science and Technology of China(中国科学技术大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

AI总结 本文提出GS-STVSR框架,通过二维高斯散射实现高效的连续时空视频超分辨率,利用高斯核的时空演变避免密集网格查询,提升处理效率和质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18032 2026-04-21 cs.CV

CFSR: Geometry-Conditioned Shadow Removal via Physical Disentanglement

CFSR:通过物理解耦实现几何条件下的阴影去除

Pan Wang, Yihao Hu, Xiujin Liu, Hang Wang

机构 * University of Science and Technology of China(科学技术大学) University of Michigan - Ann Arbor(密歇根大学安娜堡分校) Sun Yat-sen University(中山大学)

AI总结 本文提出CFSR框架,通过物理约束恢复过程解决传统阴影去除网络缺乏物理可解释性的问题,结合3D几何信息与大规模基础模型语义,有效弥合2D-3D领域差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18731 2026-04-21 cs.CL cs.AI

One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment

一个适应任何:面向个性化大语言模型对齐的元奖励建模

Hongru Cai, Yongqi Li, Tiezheng Yu, Fengbin Zhu, Wenjie Wang, Fuli Feng, Wenjie Li

机构 * The Hong Kong Polytechnic University(香港理工大学) Huawei Technologies Ltd.(华为技术有限公司) National University of Singapore(新加坡国立大学) University of Science and Technology of China(中国科学技术大学)

AI总结 本文提出元奖励建模(MRM)方法,通过元学习框架优化奖励模型初始化,提升个性化对齐的效率与鲁棒性,实验表明其在少样本场景下表现优异。

Comments Accepted by SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12554 2026-04-21 cs.CV

EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis

EmoVerse:一种基于大语言模型的emotion表示数据集用于可解释的视觉emotion分析

Yijie Guo, Dexiang Hong, Weidong Chen, Zihan She, Cheng Ye, Xiaojun Chang, Zhendong Mao

机构 * University of Science and Technology of China(中国科学技术大学)

AI总结 本文提出EmoVerse数据集,通过多层知识图谱注释实现可解释的视觉情绪分析,包含219k张图像和双注释,提供离散和连续情绪表示,结合新颖的多阶段流程和可解释模型推动高阶情绪理解。

Comments 11 pages, 7 figures. This is a preprint version of a paper submitted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26721 2026-04-21 cs.AI cs.MM

MaLoRA: Gated Modality LoRA for Key-Space Alignment in Multimodal LLM Fine-Tuning

MaLoRA:基于关键空间对齐的门控模态LoRA

Xinhan Zheng, Huyu Wu, Xueting Wang, Duo Su, Haiyun Jiang

机构 * University of Science and Technology of China(中国科学技术大学) University of Chinese Academy of Sciences(中国科学院大学) Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学) Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(巴黎-鲁维尔研究所) Rajiv Gandhi University(拉贾·甘地大学) Palmer Research Laboratories(帕勒实验室)

AI总结 本文提出MaLoRA,通过门控机制对齐多模态LLM微调中的关键空间,揭示文本偏见源于注意力键空间内在不匹配。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08878 2026-04-21 cs.SD cs.AI cs.CL eess.AS

ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling

ControlAudio:通过渐进扩散建模解决文本引导、时间指示和可理解音频生成

Yuxuan Jiang, Zehua Chen, Zeqian Ju, Yusheng Dai, Weibei Dou, Jun Zhu

机构 * Tsinghua University(清华大学) Shengshu AI(盛舒AI) University of Science and Technology of China(中国科学技术大学) Monash University(墨尔本大学)

AI总结 ControlAudio通过多任务学习和渐进扩散建模,实现文本、时间及音素特征的精细控制,提升音频生成的时序精度和语音清晰度,达到当前最优水平。

Comments Accepted at ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08252 2026-04-21 cs.IR cs.CL

ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval

ReasonEmbed:用于推理密集型文档检索的增强型文本嵌入

Jianlyu Chen, Junwei Lan, Chaofan Li, Defu Lian, Zheng Liu

机构 * University of Science and Technology of China(中国科学技术大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Beijing University of Posts and Telecommunications(北京邮电大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室) Hong Kong Polytechnic University(香港理工大学)

AI总结 本文提出ReasonEmbed,通过ReMixer合成数据、Redapter自适应学习算法和多规模backbone实现,提升推理密集型检索性能,其中ReasonEmbed-Qwen3-8B在BRIGHT基准上取得38.1的nDCG@10高分。

Comments 19 pages, 3 figures; Accepted to ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15676 2026-04-21 eess.IV cs.CV

3DGR-CT: Sparse-View CT Reconstruction with a 3D Gaussian Representation

3DGR-CT:基于3D高斯表示的稀疏视图CT重建

Yingtai Li, Xueming Fu, Han Li, Shang Zhao, Ruiyang Jin, S. Kevin Zhou

机构 * organization= School of Biomedical Engineering, Division of Life Sciences Medicine, University of Science Center for Medical Imaging, Robotics, Analytic Computing \& Learning (MIRACLE), Suzhou Institute for Advance Research, USTC , city= Suzhou , postcode= 215123 , state= Jiangsu , country= China organization= Key Laboratory of Precision organization= Key Laboratory of Intelligent Information Processing of Chinese Academy of Sciences (CAS), Institute of Computing Technology, CAS , city= Beijing , postcode= 100190 , country= China

AI总结 本文提出3DGR-CT方法,利用3D高斯表示替代隐式神经表示,通过FBP图像引导初始化和高效可微CT投影器提升稀疏视图CT重建精度与收敛速度,实现高精度实时物理模拟。

Journal ref Medical Image Analysis, Volume 103, 103585, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17325 2026-04-21 cs.CL

Align Documents to Questions: Question-Oriented Document Rewriting for Retrieval-Augmented Generation

对齐文档与问题:面向问题的文档重写以增强检索增强生成

Jiaang Li, Zhendong Mao, Quan Wang, Yuning Wan, Yongdong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Beijing University of Posts and Telecommunications(北京邮电大学)

AI总结 本文提出QREAM框架,通过风格控制重写使检索文档更符合问题需求,提升检索增强生成的准确性与效率。

Comments ACL'26 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17308 2026-04-21 cs.AI

SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents

SkillFlow:自主代理终身技能发现与演化的基准测试

Ziao Zhang, Kou Shi, Shiting Huang, Avery Nie, Yu Zeng, Yiming Zhao, Zhen Fang, Qishen Su, Haibo Qiu, Wei Yang, Qingnan Ren, Shun Zou, Wenxuan Huang, Lin Chen, Zehui Chen, Feng Zhao

机构 * University of Science and Technology of China(中国科学技术大学) University of Toronto(多伦多大学) University of Sydney(悉尼大学)

AI总结 SkillFlow通过166个任务测试自主代理能否发现、修复和维护技能库,揭示终身学习中技能演化的能力差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17278 2026-04-21 cs.CV

PestVL-Net: Enabling Multimodal Pest Learning via Fine-grained Vision-Language Interaction

PestVL-Net: 通过细粒度视觉-语言交互实现多模态害虫学习

Xueheng Li, Tao Hu, Ke Cao, Runsheng Qi, Huixin Zhang, Rui Li, Jie Zhang, Chengjun Xie

机构 * Institute of Intelligent Machines, Hefei Institutes of Physical Science, Chinese Academy of Sciences(智能机器研究所,合肥物理科学研究院,中国科学院) University of Science and Technology of China(中国科学技术大学) Zhongke Hefei Institute of Technology Innovation Engineering(中科合肥技术创新工程研究院)

AI总结 本文提出PestVL-Net框架,结合视觉与语言模型,解决细粒度害虫识别难题,通过RWKV架构和多模态大语言模型实现高效害虫学习。

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17277 2026-04-21 cs.LG cs.AI cs.ET physics.app-ph

Fully Analog Resonant Recurrent Neural Network via Metacircuit

通过元电路实现的全模拟共振递归神经网络

Zixin Zhou, Tianxi Jiang, Menglong Yang, Zhihua Feng, Qingbo He, Shiwu Zhang

机构 * Institute of Humanoid Robots, CAS Key Laboratory of Mechanical Behavior and Design of Materials, Department of Precision Machinery and Precision Instrumentation, University of Science and Technology of China(机器人研究院、材料力学行为与设计国家重点实验室、精密机械与精密仪器系、中国科学技术大学) State Key Laboratory of Mechanical System and Vibration, Shanghai Jiao Tong University(机械系统与振动国家重点实验室、上海交通大学)

AI总结 本文提出一种通过元电路架构实现的全模拟共振递归神经网络,利用共振特性实现时间信息处理,无需数字转换,展示了在触觉感知、语音识别和状态监测中的跨领域应用。

Comments 23 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11083 2026-04-21 cs.CV cs.AI

FlowCoMotion: Text-to-Motion Generation via Token-Latent Flow Modeling

FlowCoMotion:通过令牌-潜在流建模实现文本到动作生成

Dawei Guan, Di Yang, Chengjie Jin, Jiangtao Wang

机构 * School of Artificial Intelligence and Data Science, University of Science and Technology of China(人工智能与数据科学学院,中国科学技术大学) Suzhou Institute for Advanced Research, University of Science and Technology of China(苏州先进研究院,中国科学技术大学)

AI总结 本文提出FlowCoMotion框架,通过令牌-潜在耦合建模统一连续和离散动作表示,结合多视图蒸馏和时间分辨率量化,提升文本到动作生成的语义对齐与细节 fidelity。

Comments 23 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14389 2026-04-21 cs.LG cs.AI

From $\log π$ to $π$: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient Weight

从log π到π:通过双侧解耦衰减概率梯度权重来驯服发散

Xiaoliang Fu, Jiaye Lin, Yangyi Fang, Chaowen Hu, Cong Qin, Zekai Shao, Binbin Zheng, Lu Pan, Ke Zeng

机构 * Meituan(美团) Fudan University(复旦大学) Tsinghua University(清华大学) Peking University(北京大学) University of Science and Technology of China(中国科学技术大学)

AI总结 本文提出DGPO方法,通过解耦衰减机制解决RLVR中概率梯度发散问题,提升大语言模型的推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07745 2026-04-21 cs.CL cs.AI cs.LG

Parallel Test-Time Scaling for Latent Reasoning Models

潜在推理模型的并行测试时缩放

Runyang You, Yongqi Li, Meng Liu, Wenjie Wang, Liqiang Nie, Wenjie Li

机构 * Hong Kong Polytechnic University(香港理工大学) Shandong Jianzhu University(山东建筑大学) University of Science and Technology of China(中国科学技术大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

AI总结 本文提出通过改进采样和聚合方法,使潜在推理模型能有效利用并行测试时缩放,提升连续空间中的推理能力。

Comments Accepted at ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03191 2026-04-21 cs.CV

CORP: A Multi-Modal Dataset for Campus-Oriented Roadside Perception Tasks

CORP:面向校园道路感知任务的多模态数据集

Beibei Wang, Zijian Yu, Lu Zhang, Jingjing Huang, Yao Li, Haojie Ren, Yuxuan Xiao, Yuru Peng, Jianmin Ji, Yu Zhang, Yanyong Zhang

机构 * Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合国家科学中心人工智能研究院) University of Science and Technology of China(中国科学技术大学)

AI总结 本文提出CORP数据集,首个针对校园场景的多模态道路感知任务基准数据集,包含205k张图像和102k点云,用于提升校园区域的3D跟踪和实例分割性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16896 2026-04-21 q-bio.QM cs.AI

ProtoCycle: Reflective Tool-Augmented Planning for Text-Guided Protein Design

ProtoCycle:基于反思的工具增强规划用于文本引导的蛋白质设计

Yutang Ge, Guojiang Zhao, Sihang Li, Zheng Cheng, Zifeng Zhao, Hanchen Xia, Guolin Ke, Linfeng Zhang, Zhifeng Gao, Yuguang Wang

机构 * Shanghai Jiao Tong University, School of Mathematical Sciences(上海交通大学数学科学学院) DP Technology(DP技术) University of Science and Technology of China(中国科学技术大学) AI for Science Institute(AI for Science研究院)

AI总结 本文提出ProtoCycle框架,通过LLM驱动的反馈循环和轻量级工具环境,提升文本引导蛋白质设计的序列质量。

Comments 25 pages, 11 figures. Accepted to Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏