arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Texas at Austin(得克萨斯大学奥斯汀分校)

2026-08-18 至 2026-08-18 共收录 8
2608.16081 2026-08-18 cs.CV 新提交

SafeGesture: Evaluating Fine-Grained Hand Gesture Understanding in Vision-Language Models through Scenario-Conditioned Safety Interpretation

SafeGesture:通过场景条件安全解释评估视觉语言模型的细粒度手势理解能力

Taegang Kim, Saleh Afroogh, Junfeng Jiao

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Urban Information Lab, The University of Texas at Austin(德克萨斯大学奥斯汀分校城市信息实验室)

AI总结 本文提出SafeGesture基准,评估5款视觉语言模型的细粒度手势安全理解能力,发现模型存在感知与推理脱节,瓶颈为场景条件安全推理而非手势识别。

Comments 14 pages, 22 tables, 2 figures. Code and benchmark resources available at this https URL (https://github.com/The-Responsible-AI-Initiative/SafeGesture)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15962 2026-08-18 cs.CL cs.CV 新提交

SEER: Long-Context Reasoning via Selective Visual-Text Compression

SEER:基于选择性视觉-文本压缩的长上下文推理

Jiawei Xu, Zhilin Zhai, Jinrui Fang, Ruohan Xu, Mingfei Lu, Yi Zhang, Guanchu Wang, Tianlong Chen, Ying Ding

机构 * The University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of Cambridge(剑桥大学) University of Technology Sydney(悉尼科技大学) The University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 SEER是结合视觉压缩效率与文本推理精度的框架,经监督微调后在LongBench等长上下文基准测试中,准确率优于Glyph-9B、Qwen3-8B等基线模型,可提升提取精度并保留提示token节省量。

Comments COLM 2026, Third Conference on Language Modeling

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15539 2026-08-18 cs.CV 新提交

CrossView: Can Vision-Language Models Reason Across Cameras?

CrossView:视觉-语言模型能否跨相机进行推理?

Sahil Shah, S P Sharan, Harsh Goel, Manvik Pasula, Adithya Hebbalae, Minkyu Choi, Sandeep P. Chinchali

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出CrossView多相机视频问答基准,评估发现GPT-5.2等模型跨相机推理准确率低,开源模型表现更差,该基准可用于测试模型联合处理多视角的能力。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14790 2026-08-18 cs.CV 新提交

Qwen-Video-Edit: Instruction-Based Video Editing by Repurposing an Image Editing Model

Qwen-Video-Edit:通过复用图像编辑模型实现基于指令的视频编辑

Yunpeng Bai, Yossi Gandelsman, Michaël Gharbi, Qixing Huang

机构 * UT Austin(德克萨斯大学奥斯汀分校) Reve(Reve公司)

AI总结 该研究提出Qwen-Video-Edit,通过复用Qwen-Image-Edit图像编辑模型,经少量适配实现基于指令的视频编辑,证明图像编辑先验可迁移至视频编辑任务。

Comments Project Page: this https URL (https://yunpeng1998.github.io/Qwen-Video-Edit-Page;) Code: this https URL (https://github.com/yunpeng1998/Qwen-Video-Edit;) Model: this https URL (https://huggingface.co/yunpeng1998/Qwen-Video-Edit)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14707 2026-08-18 cs.AI 新提交

Semantic Uncertainty-Guided Orchestration in Hierarchical Multi-Agent Systems

分层多智能体系统中基于语义不确定性的协调策略

John Knowlton, Aritra Guha, Risto Miikkulainen

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) AT&T Chief Data Office(AT&T首席数据办公室) Cognizant AI Lab(高知特人工智能实验室)

AI总结 本文提出HASSUM框架,利用语义熵和密度估计不确定性实现多智能体自适应协调,在StrategyQA等基准上验证其可提升复杂推理任务的可靠性。

Comments 17 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21721 2026-08-18 stat.ML cs.LG 版本更新

Priors learned from legacy reconstructions inherit undetectable overconfidence

先验清洗:具有继承性、不可检测的过度自信的学习先验

Ali Siahkoohi, Sina Alemohammad

机构 * Institute for Artificial Intelligence, University of Central Florida(人工智能研究所,中央佛罗里达大学) Department of CS, University of Central Florida(计算机科学系,中央佛罗里达大学) Department of ECE, The University of Texas at Austin(电子工程系,德克萨斯大学奥斯汀分校)

AI总结 研究在贝叶斯逆问题中使用学习生成先验时因真值稀缺采用先验清洗的情况,指出其过度自信问题,通过对旧后验求平均产生旧正则化器,单个最佳档案更糟,部署时拟合档案的先验未覆盖盲子空间,建议分开数据支持置信度与继承信念。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10466 2026-08-18 cs.CV 版本更新

ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos

ExpertEdit: 从专家视频中学习技能感知的动作编辑

Arjun Somayazulu, Kristen Grauman

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 ExpertEdit通过无配对专家视频学习技能驱动的动作编辑,无需配对数据或手动指导,提升动作真实感和专家质量。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04927 2026-08-18 cs.LG cs.AI eess.SP 版本更新

Federated Self-Supervised Modulation Classification under Non-IID and Imbalanced Data

非独立同分布与不平衡数据下的联邦自监督调制分类

Usman Akram, Yiyue Chen, Haris Vikalo

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) University of Texas at Austin(德克萨斯大学奥斯汀分校) Qualcomm Technologies Inc.(高通技术公司)

AI总结 该研究针对非独立同分布与不平衡数据场景,提出FedSSL-AMC联邦自监督框架,结合因果时间膨胀CNN编码器与轻量级本地SVM,在调制分类任务上取得优于监督联邦学习基线的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏