Logo image
Query3D: LLM-Powered Open-Vocabulary Scene Segmentation with Language Embedded 3D Gaussians
Conference proceeding

Query3D: LLM-Powered Open-Vocabulary Scene Segmentation with Language Embedded 3D Gaussians

Amirhosein Chahe and Lifeng Zhou
Proceedings (IEEE Winter Conference on Applications of Computer Vision Workshops. Online), pp 961-970
28 Feb 2025

Abstract

3d gaussian splatting Computational modeling Decision making large language model Large language models Object segmentation open-vocabulary segmentation Path planning Three-dimensional displays vision language model Vocabulary Autonomous Vehicles Semantics Translation
This paper introduces a novel method for open-vocabulary 3D scene querying in autonomous driving by combining Language Embedded 3D Gaussians with Large Language Models (LLMs). We propose utilizing LLMs to generate both contextually canonical phrases and helping positive words for enhanced segmentation and scene interpretation. Our method leverages GPT-3.5 Turbo as an expert model to create a high-quality text dataset, which we then use to fine-tune smaller, more efficient LLMs for on-device deployment. Our comprehensive evaluation on the WayveScenes101 dataset demonstrates that LLM-guided segmentation significantly outperforms traditional approaches based on predefined canonical phrases. Notably, our fine-tuned smaller models achieve performance comparable to larger expert models while maintaining faster inference times. Through ablation studies, we discover that the effectiveness of helping positive words correlates with model scale, with larger models better equipped to leverage additional semantic information. This work represents a significant advancement towards more efficient, context-aware autonomous driving systems, effectively bridging 3D scene representation with high-level semantic querying while maintaining practical deployment considerations. Code and additional resources are available at https://github.com/Zhourobotics/Query-3DGS-LLM.

Metrics

Details

UN Sustainable Development Goals (SDGs)

This publication has contributed to the advancement of the following goals:

#3 Good Health and Well-Being

Source: SDGs in the Output

InCites Highlights

Data related to this publication, from InCites Benchmarking & Analytics tool:

Web of Science research areas
Computer Science, Artificial Intelligence
Computer Science, Interdisciplinary Applications
Logo image