Predicted Outputs

Not offered today. Vedika Swift plus streaming is the honest equivalent for low latency.

Coming soon
Vedika does not offer a predicted-outputs / speculative-decoding feature today. There is no parameter to enable it and no partial equivalent hiding under another name — this page exists to say so plainly and point you at what actually helps.

#Current status

Not available Coming soon. If you came here looking for a way to pre-supply an expected completion and skip regenerating unchanged tokens, that mechanism does not exist on Vedika yet.

#The honest equivalent today

What actually addresses the underlying goal — lower latency — is already live: Vedika Swift delivers roughly 1,800 tokens/sec with 1–3 second latency on Business and Enterprise plans, no prediction needed. Pair it with streaming so the first tokens render immediately. See the Model Catalog and Streaming.

#Roadmap

Not scheduled. If your use case genuinely needs predicted outputs rather than just lower latency, Vedika Swift plus streaming is the closest available substitute for now.