Singularity API
Beta is open · approvals in small batches

A lane of your own on DeepSeek V4.1 Flash

Reserved inference, booked by the hour. $0.20 per lane-hour.

Dedicated capacity for coding agents and long-context work: a 262K context window, images and tool calling on an OpenAI-compatible endpoint, available every hour of every day. Your requests never queue behind public API traffic, and your context stays cached between turns.

Your first 5 hours are on usApproved accounts start with free credits to test the lane. Limited capacity.

ONE GPU NODEYoursYOUR RESERVED SHAREOTHER LANES
Book by the hour
Reserve one or more lanes for any hour, up to a week ahead. Join a live hour for the rest of it.
Capacity that is yours
Each lane carries two concurrent requests at full speed, isolated from everyone else on the node.
Always on
The lane serves every hour of every day. Your key is permanent and goes live the moment your hour starts.
Tuned for agents
Prefix caching, speculative decoding and a 262K window keep long tool loops fast turn after turn.

Reserved lane · DeepSeek V4.1 Flash · free beta report

What the reserved lane delivered

Numbers from the free beta that preceded this launch, measured at the API with real coding-agent traffic. A reserved lane is dedicated capacity: your requests never queue behind public API traffic, your context stays cached between turns, and the speed you measure on the first day is the speed you get on the thirtieth.

Requests completed

64,381

over 30 continuous hours

Context processed

6.06 B

prompt tokens

Tokens generated

49.9 M

completion tokens

Served from cache

97.2%

of prompt tokens, billed at the cache-hit rate

Success rate

99.97%

22 failed requests, none from load

Traffic was real agentic coding work: long repository contexts, tool calls, compaction turns, and long code generations. A typical turn carried a 100K to 260K-token context of which a few thousand tokens were new, and produced a few hundred tokens of code or a tool call.

Measured at the API

Generation speed, as measured at the API

Every completed response of at least 40 tokens, timed from first token to last — everything between your client and the model included.

Tokens per second

Whole beta, all responses ≥ 40 tokens.

p10 · slowest
118
p50 · median
196
p90 · fastest
290
0100200

tok/s

The slowest tenth still ran at 118 tok/s.

Time to first token

Warm turns resume a cached context. Cold turns load one fresh.

Warm · median
1.7 s
Warm · p90
2.3 s
Cold · 100K
≈ 3.5 s
Cold · 200K
≈ 7 s
0246

seconds

  • Warm
  • Cold

About a second of the warm figure is upload: a 200K-token conversation is roughly 1 MB of JSON.

Speed under load

Per-request speed as more requests run on the lane at once. Controlled runs, 512-token answers, temperature 0.7.

0100200300
298
205
252
162
248
144
204
113
1 request4 concurrent8 concurrent16 concurrent

tok/s per request

  • Code and tool calls
  • Prose and explanations

Code outruns prose because the speculative decoder predicts it more accurately. Real agent traffic sat between the two, at a 196 tok/s median.

Specification

What the lane supports

CapabilityDetail
ModelDeepSeek V4.1 Flash, official weights. OpenAI-compatible chat completions with streaming; any OpenAI client or coding harness works by changing the base URL, key and model id.
Context window262,144 tokens. Set your harness to 262144. If a turn would overflow, max_tokens is trimmed to fit; a prompt already at the limit returns a clear error so your agent can compact and continue.
Tool callingNative function calling with parallel calls; tool calls arrive as standard tool_calls objects, streamed.
ImagesUp to 8 images per request as data URLs or public URLs (PNG, JPEG). Screenshots, diagrams, UI captures.
ReasoningOff by default for fast tool loops. Turn it on per request with reasoning_effort (low, high, xhigh, max, or a number 1–100); reasoning text arrives in the reasoning field.
Prompt cachingAutomatic and free of charge. Repeated context is served from cache and reported in usage.prompt_tokens_details.cached_tokens.
ConcurrencyTwo simultaneous requests per lane, with a short queue behind them. Book more lanes in the same hour for more parallelism. Beyond the queue you get a 429 with Retry-After.
BookingOne-hour slots, any hour of the day, up to seven days ahead. Cancel free until an hour before the slot; join a live hour for the rest of it. Your API key is permanent and is admitted only during hours you hold.
BillingPrepaid credits at $0.20 per lane-hour, bought in packs from $10 with bonus credits on larger packs. No subscription, no minimum. Disrupted hours are credited back automatically.
Data retentionZero. Prompts, images and completions are processed in memory and never stored, logged or used for training. Only billing metadata (token counts, timestamps) is kept. Read the policy.

Credits

Reserved-lane credit and Inference API credit are separate balances. The two run on different platforms, each with its own credit system, so topping up one does not fund the other. Check which account you are adding credit to before you pay.

Singularity inference API

Don’t need a whole lane? Start on the inference API.

A reserved lane is committed capacity for teams running constant traffic. If your usage is smaller or spiky, the inference API is the cheaper way in — the same platform, billed per request, with nothing to reserve and no commitment.

Models

20+

One endpoint reaches all of them.

Key

One

A single integration, not one per provider.

Caching

Semantic

Repeat work skips the round trip.

Routing

Edge

Steered to the fastest healthy region.

Testimonials

From our community

What people are saying who have already run the reserved lane.

The experience with Singularity has been really good! From API Inference (which I use daily now) to the Lane testing and now the paid lanes, everything top-notch. Speed is blazing fast, price is the minimum you can get, quality is just like the original one (GLM, DeepSeek, GPT and more) and the customer service.. 5 star!

The per hour lane system will soon explode in popularity and I'm glad to be part of it from the start!

Classified

I have good experience on Singularity API and doing an A/B test on my production application at losan.ai. I like the speed and the performance of model without sacrifice quality. I believe this is the good path because I don't have to choose between speed/quality/pricing. This is a good choice, hope you can maintain it ❤️

hiepxanh

The concepts in the beta were genuinely pretty good, and the results were impressive as well. The problems of token cost and time gates are very real, and the concept of a paid lane addresses those quite well. I had a few hiccups while testing, but it was consistent most of the time. Being able to test a model's capabilities at low or no cost before committing is another huge advantage. Definitely worth trying. 👍

XenoS1996

Singularity API is the best provider I found till date for models like DeepSeek and GLM. There is no per-M-token pricing now — you use tokens on an hourly basis without worrying about API costs, usage limits, or the 5-hour waiting period every other provider has. It serves the direct model instead of quants, which gives the best performance in intelligence, and it's blazing fast. No worrying about your data being taken for training. I'd recommend anyone to use it.

yash

Singularity API manages to provide a stable and reliable service for DeepSeek. During the beta tests, everything was very stable and completed my tasks successfully. The connection is fast and direct, and the pricing is also very affordable. Overall, I've had a great experience with it so far.

mint

Beta · pricing

One price. Book the hours you need.

A lane costs $0.20 for an hour and carries two concurrent requests; book more lanes for more parallelism. Credits come in packs from $10, with bonus credits on the larger ones. Booked hours that are disrupted are credited back automatically.

Your first 5 hours are on us.

Every approved account starts with enough free credit for five lane-hours, so you can test the lane with your own harness before buying anything. Capacity is limited while we onboard in batches.

Running ten or more hours a day? Talk to us on Discord about a committed plan.

This is a beta. Purchased credits never expire during it, and when the beta ends any unused credits are merged automatically into your account on the Singularity inference API.

Price
$0.20 per lane-hour
Concurrency
2 requests per lane
Context window
262,144 tokens
Availability
Every hour, every day
Billing
Prepaid credits, no subscription
Access
Apply, then approved by hand