MundiLABS

Computer vision · Language models · Robotics

The lab is easy. The world is not.

Mundi Labs is an independent AI practice. We take perception, language and robotics systems from a result that holds on a benchmark to one that holds in the field — and we publish what we learn along the way.

Fig. 1 · task success, deployed n = 0
measured below threshold live sample

The problem

A model that wins on the benchmark and a model that works for your users are not the same model.

The benchmark is clean, balanced and finished. The world is none of those things. Lighting changes, the camera moves, the user phrases it differently, the gripper slips, the distribution drifts, and the number that looked settled in the paper quietly stops holding.

Most teams find this out in production. They have a result they trust, a demo that impressed the room, and no instrumentation between that and the thing their customers actually touch. The fix is rarely a bigger model. It is measurement that reflects the deployed setting, and the engineering discipline that measurement makes possible.

We work on both sides of that gap, because we do not think they are separable.

Practice

Three fields, one problem

Vision, language and embodiment stopped being separate disciplines some time ago. We work across all three, and most engagements touch at least two.

Computer vision

Perception that holds under conditions the training set never contained — new sensors, new sites, new lighting, new failure modes. Detection, segmentation, depth, tracking, reconstruction and the evaluation to prove any of it moved.

  • detection & segmentation
  • depth & 3D
  • tracking
  • domain shift
  • edge deployment

Language models

Retrieval, agents and tool use taken from a prototype that demos well to a service that holds under real load, real users and a real cost ceiling — with evals that tell you the truth about whether a change helped.

  • retrieval design
  • agent architecture
  • evaluation harnesses
  • fine-tuning
  • latency & cost

Robotics & embodied systems

Policies that survive contact with hardware. Manipulation and navigation, the simulation-to- reality gap, and the unglamorous instrumentation that tells you why the robot failed at 3 a.m. rather than that it did.

  • manipulation
  • sim-to-real
  • policy learning
  • sensor fusion
  • failure analysis

Research to production

The connective tissue, and the reason to hire one team instead of two: we take a method from the literature — sometimes our own — and turn it into something you can deploy, maintain and explain.

  • method selection
  • reproduction
  • productionisation
  • benchmarking
  • handover
Method

Measure first. Then change one thing at a time.

Every engagement runs the same five stages in the same order, because skipping to the fix is how teams end up optimising something they never defined.

  1. 01

    Scope

    We agree what “working” means in your terms — not accuracy in the abstract, but the specific outcomes and failures that carry cost for your business — and what evidence would settle the question either way.

  2. 02

    Instrument

    We put measurement in place before touching a model: an evaluation set drawn from your deployed distribution rather than a public benchmark, graded against human judgement, and a harness your team can run on demand.

  3. 03

    Diagnose

    We baseline the current system, sort the failures by what they actually cost you, and tell you plainly which are worth fixing and which you should accept. Some engagements stop here, and that is a fine outcome.

  4. 04

    Build

    We work the list in priority order — architecture, data, training, retrieval, control — with every change gated on the harness, so improvement is demonstrated rather than asserted.

  5. 05

    Hand over

    You keep the harness, the datasets, the documentation and the working knowledge. The test of a good engagement is that your team can run the next one without us.

Research

We publish. That is not a side project.

Consulting firms that only consume research drift, slowly, into applying methods they no longer understand the limits of. Writing for peer review is how we keep that from happening — it forces us to state what we actually claim and to be wrong in public when we are.

It is also what you are buying. When a method comes out this month, we can tell you whether it applies to your problem, because we read it properly and often reproduced it.

Where an engagement produces something genuinely new, we will ask to publish it with you. Where it touches anything sensitive, we will not.

Engagements

Three shapes of work

Fixed scope and fixed fee wherever the work allows it. We would rather tell you the engagement is unnecessary than sell you a longer one.

Shape 01

Assessment

A short, independent read on a system you have already built. We instrument it, baseline it against your real distribution, and hand you a written verdict with a ranked list of what to do next.

2–3 weeks Fixed fee
Shape 02

Build partnership

We embed with your team and ship alongside them — writing code, reviewing designs, and raising the floor of the whole system while it goes to production.

3–6 months Embedded
Shape 03

Standing advisory

A retained line to call when the decision is expensive and reversible only at cost: architecture review, method selection, vendor diligence, or a launch that needs a second opinion.

Monthly Retainer
Contact

Tell us what you have built, and where it stops working.

We reply to every serious enquiry within two working days, usually with questions before a proposal. First conversations are free and frequently end with us saying you do not need us yet.

hello@mundilabs.ai