Terminology · Visual foundation model

Visual foundation model

Pretrained on large-scale visual data and transferable to many downstream tasks — as opposed to a model trained from scratch for one task.

Definition

A visual foundation model is pretrained on large-scale visual data and then transferred to downstream tasks. Its opposite is not a "small" model but a task-specific one, trained from scratch for a single job and rebuilt whenever the job changes.

One layer above sits the world model: it answers more than "what is this". It models space, material, light, motion and causality — how an object exists in three dimensions, how it moves, and why it moves that way. Those constraints are not in a text corpus; they are only in the physical world.

How it differs

ClassWhat it learnsWhen the task changes
Task-specific vision modelOne task's decision boundaryRetrain from scratch
Visual foundation modelTransferable visual representationsFine-tune or few-shot transfer
World modelSpace, material, light, motion, causalityOne foundation carries different tasks

Which one DaoAI World is

DaoAI World is DaoAI's in-house world-model foundation. All four product lines — ACI inspection, robot vision, SkyVision and Wemio — are built on it. The company's position: AGI is one general world model plus N vertical models, not a large language model plus skills.

The DaoAI World platform —— the foundation and the end-to-end chain

ACI · Auto Cognitive Inspection —— the inspection vertical

Robot vision —— the 3D manipulation vertical

About DaoAI —— company and technical route