Image2Gcode: Image-to-G-code Generation for Additive Manufacturing Using Diffusion-Transformer Model
Image2Gcode: 基于扩散-变换器模型的图像到G-code生成用于增材制造
Ziyue Wang, Yayati Jadhav, Peter Pak, Amir Barati Farimani
机构
*
Department of Materials Science and Engineering, Carnegie Mellon University, Pittsburgh, PA, USA(材料科学与工程系,卡内基梅隆大学,匹兹堡,PA,USA)
;
Department of Mechanical Engineering, Carnegie Mellon University, Pittsburgh, PA, USA(机械工程系,卡内基梅隆大学,匹兹堡,PA,USA)
Fast and Generalizable NeRF Architecture Selection for Satellite Scene Reconstruction
快速且通用的NeRF架构选择用于卫星场景重建
Devjyoti Chakraborty, Zaki Sukma, Rakandhiya D. Rachmanto, Kriti Ghosh, In Kee Kim, Suchendra M. Bhandarkar, Lakshmish Ramaswamy, Nancy K. O'Hare, Deepak Mishra
机构
*
School of Computing, University of Georgia(佐治亚大学计算机学院)
;
Department of Geography, University of Georgia(佐治亚大学地理系)
Reconstruction Matters: Learning Geometry-Aligned BEV Representation through 3D Gaussian Splatting
重建至关重要:通过3D高斯点云学习几何对齐的鸟瞰图表示
Yiren Lu, Xin Ye, Burhaneddin Yaman, Jingru Luo, Zhexiao Xiong, Liu Ren, Yu Yin
机构
*
Bosch Research North America \& Bosch Center for Artificial Intelligence (BCAI) Case Western Reserve University Washington University in St. Louis
机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Zhongguancun Academy(中关村学院)
;
Waytous
机构
*
Department of Computer Science, Vanderbilt University(范德比尔特大学计算机科学系)
;
Toyota InfoTech Labs(丰田信息技术实验室)
;
Department of Computer Science, University of California, Davis(加州大学戴维斯分校计算机科学系)
;
Department of Mathematics, Purdue University(普渡大学数学系)
MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
MLA:一种多感官语言-动作模型,用于机器人操作中的多模态理解和预测
Zhuoyang Liu, Jiaming Liu, Jiadong Xu, Nuowei Han, Chenyang Gu, Hao Chen, Kaichen Zhou, Renrui Zhang, Kai Chin Hsieh, Kun Wu, Zhengping Che, Jian Tang, Shanghang Zhang
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学)
;
Beijing Innovation Center of Humanoid Robotics (X-Humanoid)(北京类人机器人创新中心(X-Humanoid))
;
The Chinese University of Hong Kong (CUHK)(香港中文大学(CUHK))
Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding
无视位置,语言偏见:探测视觉语言编码器中中间层表征偏见以实现零样本语言基础空间理解
Na Min An, Inha Kang, Minhyun Lee, Hyunjung Shim
机构
*
Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
;
Samsung Electronics(三星电子)