← All open-weight models

Microsoft · Vision & multimodal

Phi-Ground-Any

Research GUI-grounding model (Phi-3-V based) that maps a natural-language instruction to on-screen coordinates — a building block for computer-use agents, not a general chat model; specs are thinly documented and it has no hosted endpoint.

Open weights · Permissive (Apache, MIT, OpenMDW) · Data checked September 6, 2026 · Not a hands-on evaluation

Model ID microsoft/Phi-Ground-Any

Parameters
Not recorded
Architecture
Not recorded
Context window
Not recordedMaximum input
Max output
Not recordedPer response
Input price
Not published
Output price
Not published
Cached input
Not published
Blended price
Not available3:1 input to output
Knowledge cutoff
Not recorded
Released
2026-05-07
Status
preview
Size band
Parameters not recorded

License and openness

Phi-Ground-Any is released under the MIT. This is a permissive license. It generally allows use, modification and redistribution, including commercial use, subject to attribution and notice requirements. License terms can change between versions; confirm the text published with the weights you download.

Capabilities

Input modalitiestext, image
Output modalitiestext

Cost at list price

InputNot published
OutputNot published
Cached inputNot published
Blended, 3:1Not available

The provider publishes no per-token price for this model. Any price you see elsewhere belongs to a third-party host, and running the weights yourself has hardware costs instead.

Estimate a workload across all models ↗

Sources

Sources checked September 6, 2026. Specifications and prices change without notice; confirm against the provider before you commit.

Other open-weight models from Microsoft

See all 23 from Microsoft ↗

Browse all open-weight models ↗