Large Models: the Cognitive Brain of PSI Agents
We aim to make large models the cognitive core of PSI agents, connecting perception, language, reasoning, and decision-making. We currently focus on two foundational capabilities: efficient multimodal inference and spatial agents in 3D environments.
- SparseVLMReduce visual tokens while retaining evidence relevant to the question.
- Proxy3DConnect compact 3D representations to a vision-language model's spatial reasoning.
More coming soon
We are exploring multi-agent evolution and desktop agents.



































