Uploaded April 2026 | Updated September 2026, 9 hours ago
Most vision systems can tell you what is in an image — WildDet3D tells you where that object sits in three dimensions, from a single photograph. Given any RGB image, it predicts 3D bounding boxes across 13K+ object categories and accepts text queries, point prompts, or 2D boxes as input — no fine-tuning required, no fixed category list, no specific hardware setup. Everything is openly available, because progress in spatial intelligence should be inspectable, reproducible, and built on by the broader research community.
Blog: allenai.org/blog/wilddet3d
Demo: huggingface.co/spaces/allenai/WildDet3D
iOS app: apps.apple.com/us/app/wilddet3d/id6760861157
Join our Discord: discord.gg/ai2
Most vision systems can tell you what is in an image — WildDet3D tells you where that object sits in three dimensions, from a single photograph. Given any RGB image, it predicts 3D bounding boxes across 13K+ object categories and accepts text queries, point prompts, or 2D boxes as input — no fine-tuning required, no fixed category list, no specific hardware setup. Everything is openly available, because progress in spatial intelligence should be inspectable, reproducible, and built on by the broader research community.
Blog: allenai.org/blog/wilddet3d
Demo: huggingface.co/spaces/allenai/WildDet3D
iOS app: apps.apple.com/us/app/wilddet3d/id6760861157
Join our Discord: discord.gg/ai2










