Joint Parsing of Cross-view Scenes with Spatio-temporal Semantic Parse Graphs

Joint Parsing of Cross-view Scenes with Spatio-temporal Semantic Parse Graphs. Qi, H., Xu, Y., Yuan, T., Wu, T., & Zhu, S. In Proceedings of The Thirty-Second AAAI Conference on Artificial Intelligence (AAAI), New Orleans, Lousiana, USA., February 2–7, pages 1–4, 2018.

Paper abstract bibtex

Cross-view video understanding is an important yet underexplored area in computer vision. In this paper, we introduce a joint parsing method that takes view-centric proposals from pre-trained computer vision models and produces spatiotemporal parse graphs that represents a coherent scene-centric understanding of cross-view scenes. Our key observations are that overlapping fields of views embed rich appearance and geometry correlations and that knowledge segments corresponding to individual vision tasks are governed by consistency constraints available in commonsense knowledge. The proposed joint parsing framework models such correlations and constraints explicitly and generates semantic parse graphs about the scene. Quantitative experiments show that scene-centric predictions in the parse graph outperform viewcentric predictions.

@InProceedings{JointParsing,
  author    = {Hang Qi and Yuanlu Xu and Tao Yuan and Tianfu Wu and Song-Chun Zhu},
  title     = {Joint Parsing of Cross-view Scenes with Spatio-temporal Semantic Parse Graphs},
  booktitle = {Proceedings of The Thirty-Second AAAI Conference on Artificial Intelligence ({AAAI}), New Orleans, Lousiana, {USA.}, February 2–7},
  year      = {2018},
  pages     = {1--4},  
  abstract  = {Cross-view video understanding is an important yet underexplored area in computer vision. In this paper, we introduce a joint parsing method that takes view-centric proposals from pre-trained computer vision models and produces spatiotemporal parse graphs that represents a coherent scene-centric understanding of cross-view scenes. Our key observations are that overlapping fields of views embed rich appearance and geometry correlations and that knowledge segments corresponding to individual vision tasks are governed by consistency constraints available in commonsense knowledge. The proposed joint parsing framework models such correlations and constraints explicitly and generates semantic parse graphs about the scene. Quantitative experiments show that scene-centric predictions in the parse graph outperform viewcentric predictions.},
  url_paper       = {https://arxiv.org/pdf/1709.05436.pdf}  
}

Downloads: 0

{"_id":"hSYbPXuGdEg8x8imC","bibbaseid":"qi-xu-yuan-wu-zhu-jointparsingofcrossviewsceneswithspatiotemporalsemanticparsegraphs-2018","downloads":0,"creationDate":"2017-11-18T19:00:17.596Z","title":"Joint Parsing of Cross-view Scenes with Spatio-temporal Semantic Parse Graphs","author_short":["Qi, H.","Xu, Y.","Yuan, T.","Wu, T.","Zhu, S."],"year":2018,"bibtype":"inproceedings","biburl":"https://tfwu.github.io/TianfuWu_BibTex.bib","bibdata":{"bibtype":"inproceedings","type":"inproceedings","author":[{"firstnames":["Hang"],"propositions":[],"lastnames":["Qi"],"suffixes":[]},{"firstnames":["Yuanlu"],"propositions":[],"lastnames":["Xu"],"suffixes":[]},{"firstnames":["Tao"],"propositions":[],"lastnames":["Yuan"],"suffixes":[]},{"firstnames":["Tianfu"],"propositions":[],"lastnames":["Wu"],"suffixes":[]},{"firstnames":["Song-Chun"],"propositions":[],"lastnames":["Zhu"],"suffixes":[]}],"title":"Joint Parsing of Cross-view Scenes with Spatio-temporal Semantic Parse Graphs","booktitle":"Proceedings of The Thirty-Second AAAI Conference on Artificial Intelligence (AAAI), New Orleans, Lousiana, USA., February 2–7","year":"2018","pages":"1–4","abstract":"Cross-view video understanding is an important yet underexplored area in computer vision. In this paper, we introduce a joint parsing method that takes view-centric proposals from pre-trained computer vision models and produces spatiotemporal parse graphs that represents a coherent scene-centric understanding of cross-view scenes. Our key observations are that overlapping fields of views embed rich appearance and geometry correlations and that knowledge segments corresponding to individual vision tasks are governed by consistency constraints available in commonsense knowledge. The proposed joint parsing framework models such correlations and constraints explicitly and generates semantic parse graphs about the scene. Quantitative experiments show that scene-centric predictions in the parse graph outperform viewcentric predictions.","url_paper":"https://arxiv.org/pdf/1709.05436.pdf","bibtex":"@InProceedings{JointParsing,\n author = {Hang Qi and Yuanlu Xu and Tao Yuan and Tianfu Wu and Song-Chun Zhu},\n title = {Joint Parsing of Cross-view Scenes with Spatio-temporal Semantic Parse Graphs},\n booktitle = {Proceedings of The Thirty-Second AAAI Conference on Artificial Intelligence ({AAAI}), New Orleans, Lousiana, {USA.}, February 2–7},\n year = {2018},\n pages = {1--4}, \n abstract = {Cross-view video understanding is an important yet underexplored area in computer vision. In this paper, we introduce a joint parsing method that takes view-centric proposals from pre-trained computer vision models and produces spatiotemporal parse graphs that represents a coherent scene-centric understanding of cross-view scenes. Our key observations are that overlapping fields of views embed rich appearance and geometry correlations and that knowledge segments corresponding to individual vision tasks are governed by consistency constraints available in commonsense knowledge. The proposed joint parsing framework models such correlations and constraints explicitly and generates semantic parse graphs about the scene. Quantitative experiments show that scene-centric predictions in the parse graph outperform viewcentric predictions.},\n url_paper = {https://arxiv.org/pdf/1709.05436.pdf} \n}\n\n\n\n","author_short":["Qi, H.","Xu, Y.","Yuan, T.","Wu, T.","Zhu, S."],"key":"JointParsing","id":"JointParsing","bibbaseid":"qi-xu-yuan-wu-zhu-jointparsingofcrossviewsceneswithspatiotemporalsemanticparsegraphs-2018","role":"author","urls":{" paper":"https://arxiv.org/pdf/1709.05436.pdf"},"downloads":0,"html":""},"search_terms":["joint","parsing","cross","view","scenes","spatio","temporal","semantic","parse","graphs","qi","xu","yuan","wu","zhu"],"keywords":[],"authorIDs":[],"dataSources":["MMF3y5eBrtyhnDQun"]}