Earticle

현재 위치 Home

텍스트 조건부 확산 모델을 이용한 3D 포인트 클라우드 생성
Text-Conditioned Diffusion Model for 3D Point Cloud Generation

첫 페이지 보기
  • 발행기관
    한국차세대컴퓨팅학회 바로가기
  • 간행물
    한국차세대컴퓨팅학회 논문지 KCI 등재 바로가기
  • 통권
    Vol.22 No.3 (2026.06)바로가기
  • 페이지
    pp.101-116
  • 저자
    장현준, 조정찬
  • 언어
    영어(ENG)
  • URL
    https://www.earticle.net/Article/A487575

원문정보

초록

영어
Current point-cloud diffusion models can generate overall shapes well, but they struggle to capture detailed semantic features and specific shape attributes described in captions. To improve this, we present a text-guided 3D point cloud generation method that uses a CLIP-conditioned diffusion model. Our approach adds natural-language information to the denoising network through token-level cross-attention and sentence-level semantic feature modulation. To prevent the model from focusing too much on less important words like articles and prepositions, we use a simple stop-word-aware attention reweighting strategy. This keeps the full token sequence but lowers the influence of stop words. We train and test our model on 3D chair and table data paired with matching natural-language descriptions. Our experiments look at overall generation quality, how well the shapes match the text, token-level attention patterns, and results from different attention setups. Quantitatively, the proposed model achieved MMD-CD scores of 6.80/6.33, COV-CD scores of 49.75/43.75, 1-NN-CD scores of 77.99/71.75, and JSD values of 9.6/13.2 on the Chair/Table categories, respectively. In the chair-category ablation study, stop-word-aware attention reweighting improved COV-CD from 37.18% to 49.75% and reduced JSD from 12.2 to 9.6. The results also show that stop-word-aware reweighting helps the model focus on key words and improves semantic coverage. Still, our model has some limits in grounding fine-grained attributes and generalizing to more object categories.
한국어
기존의 포인트 클라우드 디퓨전 모델은 전체적인 형상 생성에는 효과적이지만, 캡션에 포함된 세부적인 의미나 형상 속성을 충분히 반영하는 데 한계가 있다. 이를 해결하기 위해 본 연구는 CLIP-conditioned 디퓨전 모델을 기반으로 한 텍스트 가이드 3D 포인트 클라우드 방법을 제안한다. 구체적으로는 포인트 클라우드 확산 구조를 기반으로, 낱말 단위의 교차 주의 기법과 문장 전체 의미 기반 특징 조절 방식을 사용하여 자연어 정보를 잡음 제거 신경망에 주입한다. 또한 관사나 전치사처럼 정보량이 낮은 단어에 과도하게 주의가 집중되는 문제를 완화하기 위해, 전체 낱말 순서는 유지하면서 미리 정의된 불용어 위치의 중요도만 낮추는 경량의 불용어 인식 주의 재가중 전략을 도입한다. 우리는 의자와 탁자 범주의 3차원 형상 자료와 그에 대응하는 자연어 설명문을 사용하여 모델을 학습하고 평가한다. 실험에서는 분포 수준의 생성 품질, 문장과 형상 사이의 정성적 대응 관계, 낱말 단위 주의 동작, 그리고 주의 구성 방식에 따른 제거 실험을 분석한다. 정량적으로 제안 모델은 Chair/Table 범주에서 각각 MMD-CD 6.80/6.33, COV-CD 49.75/43.75, 1-NN-CD 77.99/71.75, JSD 9.6/13.2를 보였다. 이와 함께 Chair 범주의 ablation study에서 stop-word-aware attention reweighting은 COV-CD를 37.18%에서 49.75%로 향상시키고 JSD를 12.2에서 9.6으로 감소시켰다. 또한 불용어 인식 재가중 방식은 중요한 단어에 대한 주의 집중과 의미 반영 범위를 개선한다. 그러나 제안한 모델은 여전히 세밀한 속성 반영과 더 넓은 범주로의 일반화 측면에서 한계를 가진다.

목차

요약
Abstract
1. Introduction
2. Related Works
3. Method
3.1 Probabilistic Diffusion Model with Latent and Text Conditioning
3.2 Text-Conditioned Diffusion Network
3.3 Training Objectives
3.4 Training and Sampling
4. Experiments
4.1 Experimental Setup
4.2 Qualitative Results
4.3 Text Attention Analysis
4.4 Ablation Study
5. Limitation and Future Work
6. Conclusion
Acknowledgement
참고문헌

저자

  • 장현준 [ Hyunjun Jang | 가천대학교 인공지능학과 ]
  • 조정찬 [ Jungchan Cho | 가천대학교 인공지능학과 ] 교신저자

참고문헌

자료제공 : 네이버학술정보

간행물 정보

발행기관

  • 발행기관명
    한국차세대컴퓨팅학회 [Korean Institute of Next Generation Computing]
  • 설립연도
    2005
  • 분야
    공학>컴퓨터학
  • 소개
    본 학회는 차세대 PC 및 그 관련분야의 학술활동을 통하여 차세대 PC의 학문 및 기술발전을 도모하고 산업발전 및 국제협력 증진을 목적으로 한다.

간행물

  • 간행물명
    한국차세대컴퓨팅학회 논문지 [THE JOURNAL OF KOREAN INSTITUTE OF NEXT GENERATION COMPUTING]
  • 간기
    격월간
  • pISSN
    1975-681X
  • 수록기간
    2005~2026
  • 등재여부
    KCI 등재
  • 십진분류
    KDC 566 DDC 004

이 권호 내 다른 논문 / 한국차세대컴퓨팅학회 논문지 Vol.22 No.3

    피인용수 : 0(자료제공 : 네이버학술정보)

    함께 이용한 논문 이 논문을 다운로드한 분들이 이용한 다른 논문입니다.

      페이지 저장