Services
The engineering around the model
Naderu trains, releases and runs specialised models. The services below are that same work, done for you and handed over — not a retainer that leaves the expertise on our side of the table. Every engagement ends with artefacts you own and can rebuild without us: a versioned recipe, an evaluation suite, and a model card that states what the model does and what it does not.
What we can show you today. Naderu is early, and we would rather you heard that from us than worked it out. There are no client logos on this page and no case studies, because publishing an engagement needs a client's permission and we do not have it yet. What we can point at is the work itself: the model board is re-sourced weekly, every row carries the URL it was read from and the date it was verified, and when we get something wrong we correct it in public with our name on it. That is the standard we would hold your engagement to, and you can audit it before you talk to us.
How an engagement runs
The same four steps every time, borrowed from how we release our own models. Each step produces something you keep, so an engagement that stops early still leaves you better off than when it started.
| Step | What happens | What you keep |
|---|---|---|
| Assess | What the task actually is, what "good" means numerically, and whether a specialised model beats the general API you are already paying for. | A written assessment, including the case for not doing it. |
| Evaluate first | Build the eval before the model. Without a number agreed up front, "better" is a matter of opinion at handover — which is how these projects go wrong. | A versioned eval suite that runs on your data, in your CI. |
| Train | Fine-tune, LoRA, distill or quantize the base model that fits — chosen on the evidence, not on who we have a relationship with. | Weights, and the recipe that reproduces them: base model, method, hyperparameters, data version. |
| Hand over | Deploy where you want it and show your team how to re-run the whole pipeline. | A model card citing the exact eval suite version, and serving manifests you can
pin by name@version. |
Consulting
A scoped assessment of where a specialised model earns its keep, and where it does not. Most of the value is in the second half: the honest answer is often that a prompt and a good general model already clears the bar, and that a fine-tune would cost more than it returns. We would rather tell you that in week one than discover it in month four.
You get a written assessment naming the workload, the metric, the candidate approaches with rough cost and latency for each, and a recommendation with the reasoning attached. If the recommendation is "don't", that is the deliverable.
Model customization
The core offering, and the one built out furthest. You bring a domain and data; we turn a foundation model into one that is measurably better at your task than the general API — or we show you the numbers proving it is not, and stop.
Fine-tune, LoRA, distillation and quantization, chosen on what the evaluation says rather
than on fashion. No model is handed over without passing its eval gate,
which is the same rule our own releases live under: a model reaches models/
only after the suite passes, and the card cites the exact suite version.
You get the weights, the training recipe, the eval suite, and a model card stating provenance — including the upstream foundation model and its licence terms, because what you may do with the result depends on what you were allowed to start from.
Deployment
The model stood up where you want it: your cloud, your private cloud, on-premises, or at the edge. Serving recipes and machine-readable manifests, so the thing your application pins is a version rather than an endpoint that quietly changes underneath it.
You get the serving configuration, a manifest your consumers pin by
name@version, and a validator that fails when the deployment and the
manifest drift apart.
Operations
Models decay because the world moves, not because the weights change. Continuous evaluation against the suite built during customization, so regressions surface as a failing number rather than as a support ticket six weeks later.
You get scheduled re-evaluation, a report each cycle, and a documented path to retrain when the numbers say it is time.
Policy-tuned models
A productized form of customization for regulated, on-premises and air-gapped agents. You hand us a policy spec — blocked domains, forbidden ports, commands the agent must never run, data categories that must not leave the box — and we deliver a model with those restrictions trained into the weights, plus the adversarial evaluation that measures adherence: refusal rate on forbidden actions across direct asks, prompt injection, roleplay and encoding attacks, and over-refusal rate on allowed actions held to the same bar.
Honest scope. Fine-tuning produces a behavioural tendency, not a guarantee. Hard enforcement always belongs outside the model — sandbox, network policy, allowlisted tool schemas. What this offering sells is defence in depth and a measured number, never "no matter what". It is listed here as in development because the adversarial harness is not finished, and we will not take an engagement against it until it is.
How we choose the model
Model-agnostic by policy. We pick the best foundation model for the job and prefer open weights where they fit, because a model you can host is a model you can keep. We have no reseller relationship with any provider, and nothing on the model board is sponsored — which is checkable, since every row names its source and the date we read it.
What we will not do
- Ship a model without an evaluation. If we cannot measure it, we cannot claim it, and you should not buy it.
- Quote a number we have not measured. Including in a pitch.
- Lock you into us. The recipe and the suite are yours; the point is that your team can run them after we have gone.
- Take an engagement we do not think will work. Saying so early is cheaper for both of us than saying so late.
Starting a conversation
Email contact@naderu.com with the task you have in mind and roughly what "working" would look like. A first call is a scoping conversation, not a pitch — and if the answer is that you do not need us, that is a perfectly good outcome of it.
If you would rather look before you speak: the model board is public and refreshed weekly, its JSON is open with no key, and the weekly reports are the same analysis written up. Read a few and decide whether we think the way you want your engineers to think.