| 计算机技术 |
|
|
|
|
| 基于知识图谱引导的三维视觉问答 |
毛爱华( ),陈思宇 |
| 华南理工大学 计算机科学与工程学院,广东 广州 511400 |
|
| 3D visual question answering guided by knowledge graph |
Aihua MAO( ),Siyu CHEN |
| School of Computer Science and Engineering, South China University of Technology, Guangzhou 511400, China |
| 1 |
AZUMA D, MIYANISHI T, KURITA S, et al. ScanQA: 3D question answering for spatial scene understanding [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 19107–19117.
|
| 2 |
DELITZAS A, PARELLI M, HARS N, et al. Multi-CLIP: contrastive vision-language pre-training for question answering tasks in 3D scenes [EB/OL]. (2023–06–04)[2025–07–02]. https://arxiv.org/pdf/2306.02329.
|
| 3 |
RADFORD A, KIM J W, HALLACY C, et al. Learning transferable visual models from natural language supervision [EB/OL]. (2021–02–26)[2025–07–02]. https://arxiv.org/pdf/2103.00020.
|
| 4 |
JIN Z, HAYAT M, YANG Y, et al. Context-aware alignment and mutual masking for 3D-language pre-training [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023: 10984–10994.
|
| 5 |
CHEN D Z, CHANG A X, NIEßNER M. ScanRefer: 3D object localization in RGB-D scans using natural language [C]// Computer Vision – ECCV 2020. [S.l.]: Springer International Publishing, 2020: 202–221.
|
| 6 |
HONG Y, ZHEN H, CHEN P, et al. 3D-LLM: injecting the 3D world into large language models [EB/OL]. (2023–07–24)[2025–07–02]. https://arxiv.org/pdf/2307.12981.
|
| 7 |
LI J, LI D, SAVARESE S, et al. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models [C]// Proceedings of the International Conference on Machine Learning. [S.l.]: PMLR, 2023: 19730-19742.
|
| 8 |
ALAYRAC J B, DONAHUE J, LUC P, et al. Flamingo: a visual language model for few-shot learning [EB/OL]. (2022–11–15)[2025–07–02]. https://arxiv.org/pdf/2204.14198.
|
| 9 |
PAPINENI K, ROUKOS S, WARD T, et al. BLEU: a method for automatic evaluation of machine translation [C]// Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. Philadelphia: Association for Computational Linguistics, 2002: 311–318.
|
| 10 |
VEDANTAM R, ZITNICK C L, PARIKH D. CIDEr: consensus-based image description evaluation [C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Boston: IEEE, 2015: 4566–4575.
|
| 11 |
ZHU Z, MA X, CHEN Y, et al. 3D-VisTA: pre-trained transformer for 3D vision and text alignment [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris: IEEE, 2024: 2899–2909.
|
| 12 |
VO D M, CHEN H, SUGIMOTO A, et al. NOC-REK: novel object captioning with retrieved vocabulary from external knowledge [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 17979–17987.
|
| 13 |
SPEER R, CHIN J, HAVASI C ConceptNet 5.5: an open multilingual graph of general knowledge[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2017, 31 (1): 1
doi: 10.1609/aaai.v31i1.11164
|
| 14 |
DING Z, HAN X, NIETHAMMER M. VoteNet: a deep learning label fusion method for multi-atlas segmentation [C]// Medical Image Computing and Computer Assisted Intervention – MICCAI 2019. [S.l.]: Springer, 2019: 202–210.
|
| 15 |
PENNINGTON J, SOCHER R, MANNING C. GloVe: global vectors for word representation [C]// Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing. Doha: Association for Computational Linguistics, 2014: 1532–1543.
|
| 16 |
SCHUSTER M, PALIWAL K K Bidirectional recurrent neural networks[J]. IEEE Transactions on Signal Processing, 1997, 45 (11): 2673- 2681
doi: 10.1109/78.650093
|
| 17 |
KIPF T N, WELLING M. Semi-supervised classification with graph convolutional networks [EB/OL]. (2017–02–22)[2025–07–02]. https://arxiv.org/pdf/1609.02907.
|
| 18 |
RUMELHART D E, HINTON G E, WILLIAMS R J Learning representations by back-propagating errors[J]. Nature, 1986, 323 (6088): 533- 536
doi: 10.1038/323533a0
|
| 19 |
LIN C Y, HOVY E. Automatic evaluation of summaries using N-gram co-occurrence statistics [C]// Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology - NAACL '03. Edmonton: ACL, 2003: 71–78.
|
|
Viewed |
|
|
|
Full text
|
|
|
|
|
Abstract
|
|
|
|
|
Cited |
|
|
|
|
| |
Shared |
|
|
|
|
| |
Discussed |
|
|
|
|