Everything we sell is scoped before it starts. You should be able to read what you are buying, what it costs and what you get at the end of it without a discovery call.
What you can buy
Edge AI performance audit
The starting point for almost everyone. We profile your workload on your own hardware and hand back a written roofline analysis, a ranked list of opportunities with estimated gains, and the effort each one costs. You keep the analysis whether or not you continue.
Model-to-silicon porting
A model, a board and a number written into the statement of work before we start: this latency, this power envelope, this accuracy floor. Quantisation, delegate and runtime selection, memory planning, kernel work — whatever it takes to land the target, with the acceptance test agreed up front.
On-device LLM & VLM deployment
Language and vision-language models running locally on embedded silicon — no cloud, no per-token bill, no data leaving the device. Low-bit weight quantisation, KV-cache and context strategy, prompt and decode scheduling, and an honest measurement of what the board can and cannot sustain.
Silicon selection & architecture advisory
Before the BOM is frozen: can the cheaper part carry this workload, does the NPU actually support these operators, and what will the thermal behaviour be at the third minute rather than the third second? A short engagement that regularly saves a hardware respin.
Sovereign & private AI deployment
Modern AI running inside your own perimeter — on-premises, air-gapped, or in a jurisdiction you choose. Open-weight language and vision models sized to hardware you already own or can buy outright, with no inference leaving the building, no per-token bill, and no vendor able to deprecate the model underneath you.
AI assurance & compliance evidence
The documentation trail regulated products need for an AI component — dataset and model provenance, verification and validation evidence, accuracy and robustness argumentation, and the traceability that safety and quality auditors ask for. Prepared alongside the engineering, not retrofitted after it.
Which one do you need?
- You know it is too slow but not why → start with the performance audit. It is the cheapest way to stop guessing.
- You know exactly what has to hit what number → model-to-silicon porting, with that number written into the contract.
- You want a model on the device with no cloud in the loop → on-device LLM and VLM deployment.
- The hardware is not chosen yet → silicon selection advisory, and do it now rather than after the respin.
- Your product is regulated and the AI part has no paper trail → AI assurance and compliance evidence.
- None of the above quite fits → describe the problem and we will tell you which, or that we are the wrong people for it.
Commercials
How we contract.
Small, defined and cancellable by design — we would rather earn the next engagement than lock in this one.
Fixed-fee audit
A defined scope, a defined price and a written deliverable. Nothing renews automatically and nothing depends on you continuing. It exists so that both sides are arguing about a measured roofline instead of an estimate.
Milestone project
The acceptance criterion — a frame rate, a latency ceiling, a power envelope, a memory footprint — goes into the statement of work before the first commit. Billing follows milestones, so progress and invoices stay in step.
Advisory retainer
A standing block of senior time each month for design reviews, second opinions on a vendor toolchain, and the questions that arrive at short notice during a bring-up. Useful when the need is continuous but not full time.
IP licence & reference designs
Where we already hold an optimised kernel library, a validated model or a reference stack for your silicon, licensing it is faster and cheaper than commissioning the work. Per-product or per-unit terms, with the integration support to make it land.
Describe the workload.
What it is, what it runs on, and the number you have to hit. That is enough for us to tell you which service applies.