Try the Open Model Before You Buy the GPUs
Together AI serves a large catalogue of open-weight models over one hosted API. It sits between the two usual options: the breadth of open weights without buying the GPUs to run them. That makes it most useful before a decision rather than after one, when an organization wants to know which open model is good enough for its own content before committing hardware to it.
How it connects
The model behind an agent or a node is configuration. Together is selected per deployment or per agent.
The path this opens is worth naming because it is the reason to use it. An open-weight model evaluated on Together can later be run on the organization's own hardware through the self-hosted runtimes the platform already supports, with the same workflows and no rebuild. So an evaluation here is not throwaway work: it chooses the model that the eventual on-premises deployment runs.
It remains a hosted call, so it is not a route for material that may not leave the deployment. An evaluation is usually done on a representative subset for exactly that reason, and an engagement should say so rather than leaving someone to assume the pilot content was unrestricted.
What you can do together
- Evaluate several open-weight models on your own content without provisioning GPUs for each.
- Carry the chosen model in-house afterwards through the self-hosted runtimes, with the workflows unchanged.
- Run a broad open catalogue for general work while a self-hosted model handles restricted material.
- Change model as configuration while an evaluation is still running.
A scenario
- ScopingThe engagement agrees a representative content subset, since this is a hosted call and restricted material stays out of the evaluation.
- SetupTogether is configured as the provider in AI Intelligence Hub.
- EvaluationSeveral open models are compared on the agency's own questions, changing provider settings rather than rebuilding anything.
- DecisionOne model is good enough, and the GPU specification follows from it rather than preceding it.
- In-houseThe same model runs on the new hardware through the self-hosted runtime, and the workflows do not change.
What stays where
Together supplies the hosted catalogue
The models, the serving and the licensing stay with Together.
It is a hosted call
Restricted material stays out of an evaluation run here, and that is agreed rather than assumed.
The evaluation is reusable
The chosen open model can be brought in-house on the self-hosted runtimes.
Content governance does not move
Retrieval runs under the asking user's identity whichever provider generates the answer.
Products and solutions
Next step
See it on your own Together AI instance. We will show the connection made, the data moving and the output, then size it for your deployment.
Request a demonstration or write to sales@vidizmo.ai
sales@vidizmo.ai · vidizmo.ai/integrations/catalog/together-ai