Conferences >2021 IEEE/CVF International C...

SAT: 2D Semantics Assisted Training for 3D Visual Grounding

Download PDF
Download References
Request Permissions
Save to
Alerts

Abstract:

3D visual grounding aims at grounding a natural language description about a 3D scene, usually represented in the form of 3D point clouds, to the targeted object region. ...Show More

Metadata

Abstract:

3D visual grounding aims at grounding a natural language description about a 3D scene, usually represented in the form of 3D point clouds, to the targeted object region. Point clouds are sparse, noisy, and contain limited semantic information compared with 2D images. These inherent limitations make the 3D visual grounding problem more challenging. In this study, we propose 2D Semantics Assisted Training (SAT) that utilizes 2D image semantics in the training stage to ease point-cloud-language joint representation learning and assist 3D visual grounding. The main idea is to learn auxiliary alignments between rich, clean 2D object representations and the corresponding objects or mentioned entities in 3D scenes. SAT takes 2D object semantics, i.e., object label, image feature, and 2D geometric feature, as the extra input in training but does not require such inputs during inference. By effectively utilizing 2D semantics in training, our approach boosts the accuracy on the Nr3D dataset from 37.7% to 49.2%, which significantly surpasses the non-SAT baseline with the identical network architecture and inference input. Our approach outperforms the state of the art by large margins on multiple 3D visual grounding datasets, i.e., +10.4% absolute accuracy on Nr3D, +9.9% on Sr3D, and +5.6% on ScanRef.

Published in: 2021 IEEE/CVF International Conference on Computer Vision (ICCV)

Date of Conference: 10-17 October 2021

Date Added to IEEE Xplore: 28 February 2022

ISBN Information:

ISSN Information:

DOI: 10.1109/ICCV48922.2021.00187

Conference Location: Montreal, QC, Canada

Contents

1. Introduction

Visual grounding provides machines the ability to ground a language description to the targeted visual region. The task has received wide attention in both datasets [54], [31], [19] and methods [16], [46], [53], [50]. However, most previous visual grounding studies remain on images [54], [31], [19] and videos [57], [38], [51], which contain 2D projections of inherently 3D visual scenes. The recently proposed 3D visual grounding task [1], [4] aims to ground a natural language description about a 3D scene to the region referred to by a language query (in the form of a 3D bounding box). The 3D visual grounding task has various applications, including autonomous agents [40], [47], human-machine interaction in augmented/mixed reality [20], [22], intelligent vehicles [29], [12], and so on.

References is not available for this document.

SAT: 2D Semantics Assisted Training for 3D Visual Grounding

Abstract:

Metadata

Abstract:

ISSN Information:

1. Introduction

References

IEEE Account

Purchase Details

Profile Information

Need Help?

SAT: 2D Semantics Assisted Training for 3D Visual Grounding

Alerts

Abstract:

Metadata

Abstract:

ISSN Information:

1. Introduction

Authors

Figures

References

Citations

Keywords

Metrics

References