Terminology · Visual foundation model
Visual foundation model
Pretrained on large-scale visual data and transferable to many downstream tasks — as opposed to a model trained from scratch for one task.
Definition
A visual foundation model is pretrained on large-scale visual data and then transferred to downstream tasks. Its opposite is not a "small" model but a task-specific one, trained from scratch for a single job and rebuilt whenever the job changes.
One layer above sits the world model: it answers more than "what is this". It models space, material, light, motion and causality — how an object exists in three dimensions, how it moves, and why it moves that way. Those constraints are not in a text corpus; they are only in the physical world.
How it differs
| Class | What it learns | When the task changes |
|---|---|---|
| Task-specific vision model | One task's decision boundary | Retrain from scratch |
| Visual foundation model | Transferable visual representations | Fine-tune or few-shot transfer |
| World model | Space, material, light, motion, causality | One foundation carries different tasks |
Which one DaoAI World is
DaoAI World is DaoAI's in-house world-model foundation. All four product lines — ACI inspection, robot vision, SkyVision and Wemio — are built on it. The company's position: AGI is one general world model plus N vertical models, not a large language model plus skills.
The DaoAI World platform —— the foundation and the end-to-end chain
ACI · Auto Cognitive Inspection —— the inspection vertical
Robot vision —— the 3D manipulation vertical
About DaoAI —— company and technical route