A lane of your own on DeepSeek V4.1 Flash
Reserved inference, booked by the hour. $0.20 per lane-hour.
Dedicated capacity for coding agents and long-context work: a 262K context window, images and tool calling on an OpenAI-compatible endpoint, available every hour of every day. Your requests never queue behind public API traffic, and your context stays cached between turns.
Your first 5 hours are on usApproved accounts start with free credits to test the lane. Limited capacity.
- Book by the hour
- Reserve one or more lanes for any hour, up to a week ahead. Join a live hour for the rest of it.
- Capacity that is yours
- Each lane carries two concurrent requests at full speed, isolated from everyone else on the node.
- Always on
- The lane serves every hour of every day. Your key is permanent and goes live the moment your hour starts.
- Tuned for agents
- Prefix caching, speculative decoding and a 262K window keep long tool loops fast turn after turn.
Reserved lane · DeepSeek V4.1 Flash · free beta report
What the reserved lane delivered
Numbers from the free beta that preceded this launch, measured at the API with real coding-agent traffic. A reserved lane is dedicated capacity: your requests never queue behind public API traffic, your context stays cached between turns, and the speed you measure on the first day is the speed you get on the thirtieth.
- Requests completed
64,381
over 30 continuous hours
- Context processed
6.06 B
prompt tokens
- Tokens generated
49.9 M
completion tokens
- Served from cache
97.2%
of prompt tokens, billed at the cache-hit rate
- Success rate
99.97%
22 failed requests, none from load
Traffic was real agentic coding work: long repository contexts, tool calls, compaction turns, and long code generations. A typical turn carried a 100K to 260K-token context of which a few thousand tokens were new, and produced a few hundred tokens of code or a tool call.
Measured at the API
Generation speed, as measured at the API
Every completed response of at least 40 tokens, timed from first token to last — everything between your client and the model included.
Tokens per second
Whole beta, all responses ≥ 40 tokens.
tok/s
The slowest tenth still ran at 118 tok/s.
Time to first token
Warm turns resume a cached context. Cold turns load one fresh.
seconds
- Warm
- Cold
About a second of the warm figure is upload: a 200K-token conversation is roughly 1 MB of JSON.
Speed under load
Per-request speed as more requests run on the lane at once. Controlled runs, 512-token answers, temperature 0.7.
tok/s per request
- Code and tool calls
- Prose and explanations
Code outruns prose because the speculative decoder predicts it more accurately. Real agent traffic sat between the two, at a 196 tok/s median.
Specification
What the lane supports
| Capability | Detail |
|---|---|
| Model | DeepSeek V4.1 Flash, official weights. OpenAI-compatible chat completions with streaming; any OpenAI client or coding harness works by changing the base URL, key and model id. |
| Context window | 262,144 tokens. Set your harness to 262144. If a turn would overflow, max_tokens is trimmed to fit; a prompt already at the limit returns a clear error so your agent can compact and continue. |
| Tool calling | Native function calling with parallel calls; tool calls arrive as standard tool_calls objects, streamed. |
| Images | Up to 8 images per request as data URLs or public URLs (PNG, JPEG). Screenshots, diagrams, UI captures. |
| Reasoning | Off by default for fast tool loops. Turn it on per request with reasoning_effort (low, high, xhigh, max, or a number 1–100); reasoning text arrives in the reasoning field. |
| Prompt caching | Automatic and free of charge. Repeated context is served from cache and reported in usage.prompt_tokens_details.cached_tokens. |
| Concurrency | Two simultaneous requests per lane, with a short queue behind them. Book more lanes in the same hour for more parallelism. Beyond the queue you get a 429 with Retry-After. |
| Booking | One-hour slots, any hour of the day, up to seven days ahead. Cancel free until an hour before the slot; join a live hour for the rest of it. Your API key is permanent and is admitted only during hours you hold. |
| Billing | Prepaid credits at $0.20 per lane-hour, bought in packs from $10 with bonus credits on larger packs. No subscription, no minimum. Disrupted hours are credited back automatically. |
| Data retention | Zero. Prompts, images and completions are processed in memory and never stored, logged or used for training. Only billing metadata (token counts, timestamps) is kept. Read the policy. |
Credits
Reserved-lane credit and Inference API credit are separate balances. The two run on different platforms, each with its own credit system, so topping up one does not fund the other. Check which account you are adding credit to before you pay.
Singularity inference API
Don’t need a whole lane? Start on the inference API.
A reserved lane is committed capacity for teams running constant traffic. If your usage is smaller or spiky, the inference API is the cheaper way in — the same platform, billed per request, with nothing to reserve and no commitment.
- Models
20+
One endpoint reaches all of them.
- Key
One
A single integration, not one per provider.
- Caching
Semantic
Repeat work skips the round trip.
- Routing
Edge
Steered to the fastest healthy region.
Testimonials
From our community
What people are saying who have already run the reserved lane.
The experience with Singularity has been really good! From API Inference (which I use daily now) to the Lane testing and now the paid lanes, everything top-notch. Speed is blazing fast, price is the minimum you can get, quality is just like the original one (GLM, DeepSeek, GPT and more) and the customer service.. 5 star!
The per hour lane system will soon explode in popularity and I'm glad to be part of it from the start!
ClassifiedI have good experience on Singularity API and doing an A/B test on my production application at losan.ai. I like the speed and the performance of model without sacrifice quality. I believe this is the good path because I don't have to choose between speed/quality/pricing. This is a good choice, hope you can maintain it ❤️
hiepxanhThe concepts in the beta were genuinely pretty good, and the results were impressive as well. The problems of token cost and time gates are very real, and the concept of a paid lane addresses those quite well. I had a few hiccups while testing, but it was consistent most of the time. Being able to test a model's capabilities at low or no cost before committing is another huge advantage. Definitely worth trying. 👍
XenoS1996Singularity API is the best provider I found till date for models like DeepSeek and GLM. There is no per-M-token pricing now — you use tokens on an hourly basis without worrying about API costs, usage limits, or the 5-hour waiting period every other provider has. It serves the direct model instead of quants, which gives the best performance in intelligence, and it's blazing fast. No worrying about your data being taken for training. I'd recommend anyone to use it.
yashSingularity API manages to provide a stable and reliable service for DeepSeek. During the beta tests, everything was very stable and completed my tasks successfully. The connection is fast and direct, and the pricing is also very affordable. Overall, I've had a great experience with it so far.
mintBeta · pricing
One price. Book the hours you need.
A lane costs $0.20 for an hour and carries two concurrent requests; book more lanes for more parallelism. Credits come in packs from $10, with bonus credits on the larger ones. Booked hours that are disrupted are credited back automatically.
Your first 5 hours are on us.
Every approved account starts with enough free credit for five lane-hours, so you can test the lane with your own harness before buying anything. Capacity is limited while we onboard in batches.
Running ten or more hours a day? Talk to us on Discord about a committed plan.
This is a beta. Purchased credits never expire during it, and when the beta ends any unused credits are merged automatically into your account on the Singularity inference API.
- Price
- $0.20 per lane-hour
- Concurrency
- 2 requests per lane
- Context window
- 262,144 tokens
- Availability
- Every hour, every day
- Billing
- Prepaid credits, no subscription
- Access
- Apply, then approved by hand