Skip to content
All posts
Intelligence Systems

AI Vendor Lock-In: the Questions to Ask Before You Sign

David PackmanFounder & CEO14 min read
AI vendor lock-in, the questions to ask before you sign

AI vendor lock-in is the point at which leaving an AI supplier costs more than staying. With AI it builds up in three layers, the data you put in, what the system learns while your team uses it, and the model underneath. Most buyers check the first layer, some check the second, and almost nobody asks about the third until a model is retired and the bill for rebuilding arrives.

The questions in this post are for the meeting before you sign. That is the one point at which the answers are still negotiable, which is why the cheapest time to ask what leaves in a usable form is before the contract exists. It is also a different question from the one most ownership advice answers. What you should hold once a system is built is a list of things you can point at. This post is about how to tell, from a vendor's answers, whether you will ever be able to point at them.

None of this argues against buying AI from vendors. Almost every business will, and should. Renting software is sensible, and the risk is only ever in what accumulates inside it without anyone deciding where it should live.

What is AI vendor lock-in?

Lock-in is best understood as a cost, the cost of leaving, and it grows quietly with use. On the first day of a new AI tool, switching is cheap because nothing has been taught to it. By the second year, the team has refined its instructions, corrected hundreds of outputs and approved a way of working, and the tool has become the only place that work exists.

Ordinary software lock-in is mostly about data, and the remedies are well known. AI adds two layers that ordinary software never had. The system learns from your team as they use it, and it runs on a model that the vendor did not build, does not control and will eventually have to replace. A data export covers the first layer and neither of the others.

The competition regulator has been watching the model layer for some time. The Competition and Markets Authority's 2024 review of AI foundation models set out a principle that businesses and consumers should be able to switch between different options without being locked into one provider or ecosystem. That is a principle for how the market should work, written by a regulator concerned with competition. Whether it holds for your business depends on what you ask before you sign.

The data layer

This is the familiar one, and it is covered in detail by what you actually own when you own an intelligence system, so it needs only a short summary here.

The questions are where your data lives, what format it is stored in, and what arrives on the day you leave. The good answers are specific. The data sits in storage held in your company's name, in an open format that opens without the vendor's software, and it comes back with its structure and history intact. The warning sign is an answer that stops at "you own your data", which is nearly always true and nearly always beside the point. Owning records you cannot use elsewhere is ownership in name only.

What the system learns while you use it

The second layer is where AI lock-in genuinely differs from anything that came before, and it is the layer most contracts never mention.

Every AI system that works well gets taught. Someone rewrites the instructions until the output sounds right, a colleague corrects the draft that got a customer's name wrong, and over months the answers people approve and reject become the system's sense of what good looks like. That material is your team's judgement written down, and in most setups it accumulates inside the vendor's workspace in whatever shape the vendor chose.

The government treats this category of data as something that needs an owner from the start. Its purchasing guidance requires technology contracts to be explicit about the ownership of government data, including data created through the operation of the service. The guidance is written for public sector buyers, but the distinction it draws is exactly the one a business needs. There is the data you put in, and there is the data the service creates by running. Supplier terms often deal carefully with the first and say far less about the second.

So ask where corrections, approvals and refined instructions are stored, and whether you can read them as files. Ask whether the vendor may use any of it to train or improve its own models, and whether that is on or off by default. Ask what happens to it when you leave. A good answer describes a store you can open. A warning sign is "the system learns from your team" with no account of where that learning goes, because if it lives only inside the vendor's product, changing supplier means starting the teaching again.

This is also where the question of who owns the prompts and fine-tunes belongs. The prompts are easy to settle, because they are text and they should sit somewhere you hold. Fine-tunes are harder, and the reason is the third layer.

The model underneath

Almost no AI vendor builds its own model. It builds on one from a model provider, and model providers retire models on a schedule. That schedule reaches you whether or not your vendor mentions it.

The published policies show how much the notice varies. OpenAI commits to at least 6 months' notice before retiring a generally available model, while preview models may be retired with much shorter notice, such as 2 weeks. Google lists its stable models as available for at least 12 months after release, and once a retirement is scheduled it posts a fixed date that gives you at least 45 days to migrate. Anthropic gives at least 60 days for publicly released models. Those are the providers' commitments to their direct customers. If your vendor sits between you and the provider, your notice is whatever is left once your vendor has noticed, tested and decided what to tell you.

Fine-tunes are tied to this clock. OpenAI's own deprecations page is explicit about it. Inference on fine-tuned models will continue to be available until the base models are deprecated. On terms like these, a fine-tune can be yours in every legal sense and still stop working on a date someone else sets, and a new fine-tune on a new model has to be trained again from the data. This is one more reason memory and retrieval suit most businesses better than fine-tuning. Knowledge kept in readable pages carries across to the next model unchanged.

So ask which model the system runs on today, who decides when it changes, how much notice you will get, and whether the system could be pointed at a different model without being rebuilt. The good answer names the model, commits to a notice period and describes a layer that keeps the model replaceable. The warning sign is "we use the best model available" with nothing behind it, which tells you the vendor chooses and you find out afterwards.

We saw the value of a replaceable model with Excellerate Services' EMEA content engine. The engine sits on n8n, with frontier large language models orchestrated through a routing layer, so the model behind each regional agent can change without the workflows being rebuilt. Content production fell from 12 hours a week to about 2, and the saving does not hang on one provider keeping one model alive.

The questions, and how to read the answers

What to askA good answerThe warning sign
Where does our data live, and in what format?Storage in our company's name, in an open format we can open without you"You own your data", with no format or location
What arrives on the day we leave?A named list, with a format and a timescale"We will work with you on that at the time"
Where do our corrections and approvals end up?A store we can read and take with us"The system learns from your team", and nothing about where
Can you use our data to train your models?No, or only with our explicit consent, and off by defaultA reference to a policy page that changes without notice
Who owns the prompts and fine-tunes?We do, and the prompts are files we holdOwnership on paper of something we cannot read
Which model does it run on, and who decides when that changes?A named model, a notice period, and our sign-off on changes"We always use the best model available"
Could we point it at a different model?Yes, through a layer built for it, followed by testingA rebuild, or a pause while they think about it

What an answer says matters, and so does how quickly and specifically it arrives. A vendor that has designed for portability answers these questions in a sentence, because somebody has already decided. A vendor that has not will usually answer honestly but slowly, and a long pause on the model questions in particular is worth taking seriously. That is rarely bad faith, and far more often a sign that the question has never come up, which tells you the arrangement is whatever happens to be convenient for the vendor at the time.

Ask the same questions of a partner who builds for you as of a product you buy. The difference is that a build partner can usually say yes to all of them, because the choices are being made for you rather than for a product with thousands of customers. Lock-in is a choice the builder makes, and the questions above are how you find out which choice is being made.

What a contract can do, and what it cannot

A contract records an arrangement, and the conditions that make the arrangement worth anything come from how the system is built. That point is made at more length in what you actually own, and for AI it has a specific edge, because a clause can oblige a vendor to return your data but cannot keep a retired model running.

That said, three clauses are worth asking for specifically because they cover the layers that standard terms tend to miss. The first is ownership of the data created while the system runs, written so that it plainly includes corrections, approvals and refined instructions as well as the data you supply. The second is notice of any change to the underlying model, with a period long enough to test the replacement before it goes live. The third is an exit clause that names the format, the contents and the timescale of what comes back, so that leaving is a process rather than a negotiation.

Beyond those three, the strongest protection is the design itself. If the knowledge lives in files you hold, the learning lands somewhere you can read and the model sits behind a layer you can repoint, the contract has very little left to do. That is the principle behind owning your intelligence, and it is the standard we hold our own builds to.

Practical takeaways

  1. Ask about all three layers. A good data answer tells you nothing about the learning or the model, and those are the two layers that make AI lock-in different from the software lock-in you already know how to manage.
  2. Ask where the teaching goes. Corrections, approvals and refined instructions are your team's judgement, and they should land somewhere you can read and keep.
  3. Get the model named. Ask which model the system runs on, who decides when it changes and how much notice you get, because every model provider retires models on its own schedule.
  4. Be wary of relying on fine-tunes. With at least one major provider, a fine-tune runs only as long as its base model does, so knowledge kept in readable pages is the safer place for anything the business needs to keep.
  5. Judge the speed of the answer. A vendor that has designed for portability answers in a sentence. A long pause usually means nobody has decided.

Frequently asked questions

What is AI vendor lock-in?

It is the point at which leaving an AI supplier costs more than staying, even when staying is no longer the right call. With ordinary software the cost is mostly the data. With AI it sits in three layers. There is the data you put in, there is what the system learns while your team uses it, meaning the corrections, approvals and refined instructions, and there is the model underneath, which the vendor chooses and which its own provider will eventually retire. A business can have a clean data export and still be locked in at the other two.

What questions should we ask an AI vendor?

Ask one set of questions for each layer. For the data, ask where it lives, what format it comes back in and what arrives on the day you leave. For what the system learns, ask where your team's corrections and approvals are stored, whether you can read them, and whether the vendor may use them to train anything. For the model, ask which one it runs on, who decides when that changes, how much notice reaches you, and whether you could point the system at a different model without rebuilding it. Listen for specific answers, because a vague one usually means nobody has decided.

Who owns the prompts and fine-tunes we pay for?

On paper, usually you, and that is the less important half of the question. Vendor terms commonly assign you your inputs and outputs, but a prompt you own and cannot read is of little use, and a fine-tune only runs on the model it was trained from. One major provider states that inference on fine-tuned models continues only until the base models are deprecated, so a fine-tune you own outright still has an expiry date set by somebody else. Ask for the prompts as readable files you hold, and treat any fine-tune as something you may have to pay to recreate.

Can we switch AI models without starting again?

Yes, if the system was built for it, and it is worth checking before you sign. The model should sit behind a layer that lets you point the same workflows at a different model, and what the business knows should live in readable pages the system retrieves from, not inside one model's trained weights. Built that way, a model change is a configuration change followed by testing. Built the other way, every refined prompt and every fine-tune has to be redone. Model providers retire models routinely, so this happens to everybody eventually.

Can a contract prevent AI vendor lock-in?

Only partly. A contract can require a vendor to hand things back, give notice of changes and help with an exit, but no clause makes a proprietary format readable or keeps a retired model running. Only the way the system is built can do that. The clauses worth having are ownership of the data created while the system runs, not just the data you supply, notice of any change to the underlying model, and an exit that names the format and the timescale. Ask for those three, and judge the vendor more by how the system is built than by what the contract promises.


Related Articles