AI strategy & discovery
We inventory the candidate use cases, test feasibility, work out the cost and return, then recommend the process and ownership changes that go with it. Most projects lose half their candidate list at this step.
Milisoft offers four services: strategy consulting, custom development, model engineering and operations support. Buy one piece to cover a gap in your own team, or hand us the whole path — from use-case discovery through to long-term on-call cover after launch.
Early on what's missing is judgement; in the middle it's engineering; after launch it's someone watching the system. Each of the four below can be bought on its own, or strung together into one delivery path.
We inventory the candidate use cases, test feasibility, work out the cost and return, then recommend the process and ownership changes that go with it. Most projects lose half their candidate list at this step.
Agents built the way software gets built: eval set first, code review, canary release. Source, deployment scripts and an operations runbook are handed over, so your team can carry it.
Model selection and benchmarking, instruction tuning, distillation and quantisation, inference performance and cost tuning. Sovereign hardware and software, and fully offline installs, are both supported.
Monitoring and alerting, release regression, cost reviews, incident response and quarterly optimisation reviews. Three SLA tiers set the response times and the hours we cover.
We start with business interviews and an inventory of the systems you already run, then put every candidate use case on the same scorecard: can the data be reached, is the permission boundary clear, who covers a wrong answer, and do the hours saved convert into money. Half the list usually gets cut. What survives is worth a development budget.
Requirements, interface design, eval set first, code review, canary release — we build agents the way software gets built, not by tuning a prompt until it demos well. A project team is typically one architect, two to four engineers and one evaluation engineer, sized up or down by stage.
We benchmark candidate models side by side on your own business samples before deciding whether to fine-tune at all — training first is the wrong order. After that the work is fitting the model into the GPU memory and the budget you actually have: distillation, quantisation, batching and caching are tuned together until P95 latency and cost per call both land in a range you can accept.
After delivery we keep watching the system: call volume, latency, failure rate and token cost all carry alert thresholds. When a model provider ships an update, the knowledge base is restructured or a business definition changes, the eval regression runs before anything is let through — so an answer that was right yesterday doesn't quietly go wrong today.
An internal pilot and a system your customers depend on need very different levels of cover. The three tiers below differ mainly in service hours, response speed and how often we tune. Sign per system — there is no need to put the whole company on one tier.
For internal pilots, departmental tools and non-critical business systems
For systems already in production with steady daily use, at department or cross-department scale
For customer-facing systems, systems that affect revenue, and systems under regulatory requirements
| Capability | Standard | Professional | Flagship |
|---|---|---|---|
| Service hours | Business days 9:00–18:00 | Business days 9:00–21:00 | 7×24, including public holidays |
| First response to alerts | Within 4 hours | Within 1 hour | Within 15 minutes |
| Recovery target | 2 business days | 8 hours | 4 hours |
| Named technical contact | — | ● | ● plus a backup contact |
| Optimisation review | — | ● quarterly | ● monthly |
| Eval regression frequency | Twice a year | Quarterly | Monthly, and before every change |
| Version upgrade support | Remote guidance | Remote implementation | On-site implementation |
| Emergency change window | — | ● requested 1 business day ahead | ● requested at any time |
The table above describes the service tiers in general terms. The specific SLA metrics, how they are measured, where responsibility sits and which exceptions apply are governed by the service contract signed by both parties.
Consulting is priced per project, with the scope and deliverables listed line by line in the contract. Custom development is quoted in person-days and accepted stage by stage — if a stage doesn't pass, the next one doesn't start. Operations support is an annual subscription; the price depends on the tier you choose and how many systems we cover.
Compute or API charges for model inference are normally paid by you directly. We give an estimated range in the proposal and reconcile it by department once the system is live, so nobody discovers an overrun at the end of the month.
Yes. Consulting is delivered as a standalone service. You end up with a ranked use-case list, a feasibility report, and a budget and schedule recommendation — build from it yourself or hand it to another team. There is no tie-in clause.
Some clients finish the consulting and conclude that the right next step is data governance, with AI deferred. That is a valid outcome too.
Intellectual property in work built specifically for you belongs to you. Source, deployment scripts and architecture documents are delivered with the project, and we run a code walkthrough so your engineers can read it and change it.
Milisoft's own platform modules are provided under licence. The scope of that licence, what you may modify and how upgrades work are agreed separately in the contract — you will not end up with a delivery you cannot touch.
Yes. The default is remote work with people on site at the moments that matter: requirements clarification, production integration testing, cut-over and training handover. For projects involving isolated networks, data that cannot leave the premises, or sovereign environments, engineers can be on site for the whole engagement.
Headcount, duration and cost for on-site work are agreed separately. The Flagship tier includes on-site implementation for major version upgrades by default.
Public cloud, private cloud, your own data centre, and fully offline internal networks. The system ships containerised and runs on Kubernetes, with a single-machine install supported for smaller deployments.
For sovereign environments we support domestically-controlled chips and operating systems. Specific models need an environment check before the project starts; we provide a checklist and the minimum resource requirements.
We fix it first and settle the cost afterwards. If production is unavailable, then whatever tier you are on, raise a ticket or write to contact@milisoft.com to escalate. The on-call engineer is alerted and starts containing the problem straight away.
The work is then billed on actual effort, with a written post-incident report covering root cause and the fixes. If the same class of problem keeps recurring, we will recommend a change of tier or of architecture rather than keep charging per incident.
Describe the one thing you most want automated. In 60 minutes we can usually tell whether it is a consulting, an engineering or an operations problem, and we follow up with a clear first step. Pre-sales and technical support both reach us at contact@milisoft.com.