APIPIE LABS

Building on frontier and open-weight models

Everything between the model and the user.

We design, build and run the AI, cloud, security and data systems behind modern products — from a single model endpoint to a platform your whole company depends on.

Traffic · live
routed rerouted
Routing core Live
RouteResolvedp95State
support.replyprovider‑03 lg412msok
doc.extractprovider‑07 md288msok
code.reviewprovider‑01 lg905msok
ticket.triageprovider‑11 sm154msok
search.rerankprovider‑04 sm96msok
Illustrative — eleven providers, failover mid‑flight
2.4MRequests served daily
99.98%Platform uptime
41Models in production
+38msAdded p50 latency
31%Median spend cut

Who is behind this

Most of this job is unglamorous. A provider rate-limits at four in the morning, a schema quietly changes shape, a permission nobody scoped turns out to matter. We are the people who get paged for that.

Signed by the team
Kuala Lumpur

There are eight of us. You will speak to the engineer who does the work — on the first call and on every call after it. Nobody here is paid a commission to close you, which is why saying no to a third of enquiries costs us nothing.

What we do

Six things, built to work together.

Most vendors sell one layer and hand you the integration problem. We build the whole path — model, infrastructure, security, data — so nothing falls between the seams.

Artificial Intelligence01

Agents that finish work

Reading the ticket, checking the system, drafting the reply, escalating where being wrong is expensive.

  • Retrieval over your own documents
  • Evaluation suites and accuracy thresholds
  • Fine-tuning and open-weight deployment
  • Human-in-the-loop escalation
Cloud02

Survives a bad Tuesday

Autoscaling, multi-region failover and cost controls that stop a runaway job becoming a five-figure invoice.

  • AWS · GCP · Azure · bare metal
  • Kubernetes and serverless GPU
  • Infrastructure as code, reviewed like code
  • Spend alerting and circuit breakers
Cybersecurity03

The surface AI just added

Prompt injection, exfiltration through a model, over-permissioned agents — new attack surface most teams have not mapped yet.

  • AI-specific threat modelling
  • Secrets, key rotation and scoped access
  • Audit trails an auditor can read
  • PDPA and GDPR posture
Data04

An index that stays honest

The pipelines, versioning and freshness checks that stop a model confidently quoting last quarter.

  • Pipelines, warehousing and lineage
  • Vector stores and hybrid retrieval
  • Quality monitoring and drift alerts
  • Real-time and batch, same contract
Software05

The application around it

Web, mobile and internal tools built by the same team that built the inference layer, so the seams are ours to answer for.

  • Full-stack product engineering
  • API design and versioning
  • Design systems and front-end
  • Handover with a tested runbook
Products06

Tools we built for ourselves

A routing gateway, an evaluation harness and a monitoring console — available to you directly.

  • Unified model gateway
  • Evaluation and regression harness
  • Cost attribution console
  • Self-host or managed

Deployed where your users are

Eleven regions, automatic failover.

Traffic reroutes the moment a provider degrades — without a line of change in your application.

How we work

Four stages. Stop after any of them.

Each stage produces an artefact in your own repository. You can stop at any point and what exists still runs.

01 · Trace

Read a week of real traffic

We instrument what you already run. Most of what happens next comes out of that log, not a workshop.

02 · Route

One workload moves

Behind the gateway with your fallback list and thresholds. Nothing user-facing changes while the numbers come in.

03 · Harden

Bounds and ceilings

Schema validation, bounded retries, spend ceilings and the alerting that tells you before a customer does.

04 · Hand over

It goes in your repo

Config, evaluation set and runbook. Whether we stay on is a decision you make with the numbers in front of you.

“The handover is the part clients tell us they were not expecting. Everything is in your repository, written so the person who joins next year can read it without us in the room.”

— Placeholder attribution · swap for a real engineer

The people

You will know who is doing it.

There are eight of us, in one room in Kuala Lumpur. On your engagement you will deal with four — and you will deal with them directly.

No account manager sits in between. Nobody here earns a commission for keeping you on the phone, which is why we can afford to tell a third of enquiries that we are the wrong people for the job.

No hand-off after the sale. The people you meet before the project are the people responsible for shipping it.

Written as roles rather than job titles — names and photographs replace them before launch.

Next step

Point one workload at the gateway.

A sandbox key takes minutes and a fortnight of shadow traffic tells you more than any proposal will. You keep the routing report either way.

Kuala Lumpur · the whole team, one room.

Services

Everything between the model and the user.

Anyone can call a model. What breaks in production is the layer around it — the provider that rate-limits at 4am, the schema that drifts, the permission nobody scoped. That layer is what we build.

Artificial Intelligence

Agents that finish work rather than demo it — reading the ticket, checking the system, drafting the reply, escalating where being wrong is expensive.

  • Retrieval over your own documents
  • Evaluation suites and accuracy thresholds
  • Fine-tuning and open-weight deployment
  • Human-in-the-loop escalation

Cloud

Infrastructure that survives a bad Tuesday. Autoscaling, multi-region failover and cost controls that stop a runaway job becoming a five-figure invoice.

  • AWS · GCP · Azure · bare metal
  • Kubernetes and serverless GPU
  • Infrastructure as code, reviewed like code
  • Spend alerting and circuit breakers

Cybersecurity

Prompt injection, exfiltration through a model, over-permissioned agents — the attack surface AI adds is new, and most teams have not mapped it yet.

  • AI-specific threat modelling
  • Secrets, key rotation and scoped access
  • Audit trails an auditor can read
  • PDPA and GDPR posture

Data

A retrieval layer is only as honest as the index behind it. We build the pipelines, the versioning and the freshness checks that stop a model confidently quoting last quarter.

  • Pipelines, warehousing and lineage
  • Vector stores and hybrid retrieval
  • Quality monitoring and drift alerts
  • Real-time and batch, same contract

Software

The application around the model. Web, mobile and internal tools built by the same team that built the inference layer, so the seams are ours to answer for.

  • Full-stack product engineering
  • API design and versioning
  • Design systems and front-end
  • Handover with a tested runbook

Products

Our own tools, available to you directly — a routing gateway, an evaluation harness and a monitoring console we built because we needed them ourselves.

  • Unified model gateway
  • Evaluation and regression harness
  • Cost attribution console
  • Self-host or managed

Effort and price

Indicative bands.

Ranges from engagements we have actually closed, published so a finance conversation can happen before an engineering one. The binding figure arrives at the end of the discovery week.

ServiceTypical buildIndicative bandMonthly runMain risk
Artificial Intelligence5–8 weeksUSD 28k–60kUSD 3k–8kNo agreed definition of a correct answer
Cloud2–6 weeksUSD 14k–40kUSD 1k–4kMigration scope creeping past the brief
Cybersecurity3–5 weeksUSD 12k–32kIncluded in runFindings nobody is resourced to fix
Data3–8 weeksUSD 15k–45kUSD 1k–4kTurning into a full data programme
Software6–12 weeksUSD 30k–90kUSD 2k–6kRequirements arriving after the build starts
ProductsSame weekNo build feeUSD 500–5kPaying for a console nobody opens

Bands assume one system, not a per-seat licence, and they exclude the model spend itself — that goes on your own provider account, so the number you see there is the number they charged. Platform pricing is separate and starts where the routing saves more than it costs.

“We publish these because the alternative is three meetings before anybody says a number out loud. If your budget sits under the bottom of a band, tell us on the first call — sometimes the honest answer is a smaller piece of work, and sometimes it is nothing at all.”

— Placeholder attribution · swap for a real engineer

Questions

Asked on almost every first call.

Self-hosted or managed — which should we pick?

Self-host if a payload leaving your network is a compliance problem, or if you already run Kubernetes and would rather own the uptime. Managed if you want the routing table maintained for you as provider prices move. The code is identical; only who runs it changes.

What happens when a provider has an outage?

The request moves to the next entry in your fallback list mid-flight, and the caller sees a normal response with a different resolved model in the trace. You get an alert; your users get nothing unusual.

If every provider on a route is down, the gateway returns a structured error rather than a timeout, so your application can degrade on purpose instead of hanging.

Can we bring our own provider keys?

Yes, and most teams do. Your keys stay in your secret manager and are used at request time. You keep the direct billing relationship and we never hold a credential that can spend your money.

How do you decide which model wins a route?

Three signals, weighted by you: measured accuracy on your own evaluation set, p95 latency against the budget for that route, and cost per completed task including retries. Any model can be pinned by hand, and pinning is logged so nobody quietly overrides the data.

Do you take on work outside the platform?

Yes — cloud, security, data and software engagements that stand on their own. We take roughly six to eight a year and turn down anything where we would be the third vendor in a room with no technical owner.

APIPIE Gateway · control plane · self-host or managed

One control plane over every provider.

Requests are scored on cost, latency and measured quality, then routed to whichever provider wins that task. When one degrades, traffic moves mid-flight and your callers never notice.

or

Providers
11
Regions
11
Added p50
+38ms
Failover
Mid‑flight
Keys
Yours
Trace
Every call
Fig. 1 Eleven providers behind one key. Which one answered is written into the trace, not discovered later in a support ticket.

§1  Architecture

One request, four layers.

They are not four features you choose between. Every call passes through all four in order, and the fourth writes down what the first three did. Below is one support ticket making the trip, with the line each layer left behind.

  1. L1

    Route

    Per-task scoring across eleven providers, with any model pinnable by hand. The scoring set is yours and lives in your repository.

    resolved → provider‑03 · lg
    score 0.91 — beat provider‑07 on cost
  2. L2

    Cache

    Semantic cache with a similarity threshold you set per route. On support workloads this alone removes a third of the spend.

    miss  similarity 0.71
    under your 0.86 floor, so it went to the model
  3. L3

    Guard

    Schema validation, bounded retries and a confidence floor before anything is returned to a caller.

    schema ok  confidence 0.93
    cleared the 0.90 floor — sent unattended
  4. L4

    Trace

    Input, resolved model, tokens, latency, cost and cache status logged on every single call. Exportable.

    1,284 tok  412ms  $0.0031
    written before the reply left the building

One ticket, one route, an ordinary Tuesday. Yours will read differently — the point is that it reads at all.

§2  How a route resolves

Three signals, weighted by you.

Measured accuracy on your own evaluation set, p95 latency against the budget for that route, and cost per completed task including retries. Move the weights. The winner changes, and that is the entire decision.

50
30
20
CandidateAccuracyp95Cost / taskWeighted
Fig. 2 Five candidates on one route, scored live. Illustrative — on your account these come from your evaluation set and your invoices.

§3  Inside the platform

Three views you get on day one.

Not a slide deck — the consoles your team logs into, running in your own infrastructure, over hardware you can walk up to.

Routing coreWhich model answered each task, and the score that decided it.
Capacity & quotaHeadroom per provider, per route. Over 85% is one bad hour from a queue.
Cost attributionSpend by team, task and cache status — cache hits already deducted.

§4  Honestly

Three ways to do this. We are one of them.

If your situation fits one of the other two columns, we would rather you knew now than found out in month three.

Straight to one provider

Minutes to your first call. Swapping later means a code change and a redeploy, failover is whatever retry logic you write yourself, and cost arrives as one invoice per provider with no idea which team spent it.

Right when you have one provider and one workload, and you would rather ship than architect.

Where we fit

Behind the gateway

Minutes to your first call with a sandbox key. Swapping a provider is a config change with no deploy, failover happens mid-flight, and cost is attributed per route, per task and per team. Runs in your infrastructure or ours.

Right when you run several providers and somebody is accountable for the spend.

Build it yourself

Weeks, not minutes. Everything above is achievable and some teams should absolutely do it — you own the routing table, the failover semantics and the trace format, and you also own them at four in the morning.

Right when you have a platform team with the time, and routing is close enough to your product to be worth owning.

Comparison reconstructed for this mockup — check it against the live page before publication.

§5  Getting a key

Try it against your own traffic.

Point a copy of production traffic at the gateway for a fortnight. Nothing changes for your users, and you keep the routing report either way — including the version where the report says you do not need us.

  1. Same day Sandbox key, and a route of your choosing.
  2. Day two Shadow traffic starts. Your production path is untouched.
  3. Fortnight The report: what each provider actually cost you, and where the routing would have moved.

Fig. 3 A fortnight of shadow traffic, watched by people who have done it before.

Contact

Replies within one working day.

Thirty minutes, then a written scope. Bring one process that is eating your team's week. We will tell you on the call whether it is a good candidate — and we say no to roughly a third of what we are asked about.

Write the way you would write to a colleague. One paragraph is plenty — we do not need a brief, and there is no form long enough to replace a conversation.

The room your message lands in. Somebody here opens it each morning.

Mockup form — nothing is sent from this page.

Process

Timelines you can plan around.

StageWhenWhat it is
First replyWithin one working dayFrom an engineer who has read your message
Technical callWithin a weekForty minutes, with technical requirements
Sandbox keySame day as the callFor testing staging traffic
Discovery weekTwo to five weeks outFor engagements only
Binding quoteEnd of that weekFixed scope, fee and exclusions