Services

Five ways to engage AI ThinkLab: performance audit, model-to-silicon porting, on-device LLM and VLM deployment, silicon selection advisory, and AI assurance evidence.

Everything we sell is scoped before it starts. You should be able to read what you are buying, what it costs and what you get at the end of it without a discovery call.

What you can buy

Edge AI performance audit

The starting point for almost everyone. We profile your workload on your own hardware and hand back a written roofline analysis, a ranked list of opportunities with estimated gains, and the effort each one costs. You keep the analysis whether or not you continue.

Fixed fee · 2–3 weeks · no commitment beyond it

Model-to-silicon porting

A model, a board and a number written into the statement of work before we start: this latency, this power envelope, this accuracy floor. Quantisation, delegate and runtime selection, memory planning, kernel work — whatever it takes to land the target, with the acceptance test agreed up front.

Milestone-billed · target written into the SoW

On-device LLM & VLM deployment

Language and vision-language models running locally on embedded silicon — no cloud, no per-token bill, no data leaving the device. Low-bit weight quantisation, KV-cache and context strategy, prompt and decode scheduling, and an honest measurement of what the board can and cannot sustain.

Jetson · Snapdragon · RK3588 · i.MX · x86 edge

Silicon selection & architecture advisory

Before the BOM is frozen: can the cheaper part carry this workload, does the NPU actually support these operators, and what will the thermal behaviour be at the third minute rather than the third second? A short engagement that regularly saves a hardware respin.

Pre-BOM · 1–2 weeks · design-review format

Sovereign & private AI deployment

Modern AI running inside your own perimeter — on-premises, air-gapped, or in a jurisdiction you choose. Open-weight language and vision models sized to hardware you already own or can buy outright, with no inference leaving the building, no per-token bill, and no vendor able to deprecate the model underneath you.

On-premises · air-gapped · data-residency constrained

AI assurance & compliance evidence

The documentation trail regulated products need for an AI component — dataset and model provenance, verification and validation evidence, accuracy and robustness argumentation, and the traceability that safety and quality auditors ask for. Prepared alongside the engineering, not retrofitted after it.

ISO 26262 · IEC 62304 · EU AI Act readiness

Which one do you need?

  • You know it is too slow but not why → start with the performance audit. It is the cheapest way to stop guessing.
  • You know exactly what has to hit what number → model-to-silicon porting, with that number written into the contract.
  • You want a model on the device with no cloud in the loop → on-device LLM and VLM deployment.
  • The hardware is not chosen yet → silicon selection advisory, and do it now rather than after the respin.
  • Your product is regulated and the AI part has no paper trail → AI assurance and compliance evidence.
  • None of the above quite fits → describe the problem and we will tell you which, or that we are the wrong people for it.

Commercials

How we contract.

Small, defined and cancellable by design — we would rather earn the next engagement than lock in this one.

Fixed-fee audit

A defined scope, a defined price and a written deliverable. Nothing renews automatically and nothing depends on you continuing. It exists so that both sides are arguing about a measured roofline instead of an estimate.

Priced before we start

Milestone project

The acceptance criterion — a frame rate, a latency ceiling, a power envelope, a memory footprint — goes into the statement of work before the first commit. Billing follows milestones, so progress and invoices stay in step.

Outcome written into the SoW

Advisory retainer

A standing block of senior time each month for design reviews, second opinions on a vendor toolchain, and the questions that arrive at short notice during a bring-up. Useful when the need is continuous but not full time.

Monthly · capped hours

IP licence & reference designs

Where we already hold an optimised kernel library, a validated model or a reference stack for your silicon, licensing it is faster and cheaper than commissioning the work. Per-product or per-unit terms, with the integration support to make it land.

Perpetual or per-unit terms

Describe the workload.

What it is, what it runs on, and the number you have to hit. That is enough for us to tell you which service applies.