All resources
Concepts & fundamentals

What "AI-Native" Actually Means, and How to Test the Claim

AI-powered means AI is inside your product. AI-native means AI is one of its users. Five tests that separate the two, a checklist you can run against any vendor, and an honest scoring of our own form backend against it.

The ShipMyForm team

· 9 min read

"AI-powered" describes where the AI sits: inside the product, doing some of the work. "AI-native" describes who the product is for: an AI agent is one of its users, with the same standing as a person. The two are independent, and most products claiming the second have only built the first. Below are five tests that tell them apart, and our own form backend scored against them — including the two it does not pass.

Full disclosure: ShipMyForm is our product, so the last section is self-interested. The five tests are the useful part, and they are written so you can run them against anyone, including us.

The phrase has stopped meaning anything

A year of launch posts has reduced "AI-native" to "we added a chat box". That is a shame, because there is a real distinction underneath it and it is becoming the one that matters.

Here is the version worth keeping:

  • AI-powered is about mechanism. Something inside the product is a model. A spam filter that classifies text, a summariser, a generated reply. The user is still a person; the AI is a feature they benefit from.
  • AI-native is about audience. One of the product's users is a program that reasons. It arrives with a goal rather than a click path, needs to be told what it may do, and will be the thing deciding whether to use you at all.

A product can be either, both, or neither. Our spam filtering is AI-powered and that is not what this post is about. A model in your pipeline says nothing about whether an agent can use your product.

The reason the distinction is becoming load-bearing: increasingly, the agent chooses the vendor. When someone tells Claude Code or Cursor to "add a contact form to this site", the thing evaluating the options is not reading your pricing page for the feel of it. It is looking for something it can actually finish the task with. Being pleasant for humans and impossible for agents is a new way to lose a customer you never knew you had.

Test 1: Does the machine interface cover the whole product?

The common failure is a token integration. Three tools (create a resource, read a status), and for anything else the agent is handed a dashboard link and the human takes over. That is a machine-readable brochure, not a machine-usable product.

The test: take the things a human can do in the UI and ask what fraction an agent can do. Not "is there an API": there is always an API. Can it do the whole job, including the unglamorous parts: billing, keys, configuration, deletion.

How to run it: ask the vendor's agent interface to list its own tools, then compare that list against the navigation in their dashboard. The gaps are the answer.

Test 2: Can an agent start from nothing?

This is where most products fail, and they fail before the first API call. Signing up is a web form. Then a CAPTCHA. Then a credit card. Then "click the link in your email". Every one of those is a wall to a program, and the agent's only move is to stop and ask a human to go and do it, which breaks the task in half.

The test: from no account at all, can an agent get to a working integration without a person opening a browser tab?

Note that this is not an argument for letting agents sign up unsupervised. Approval matters, and an agent that can create accounts and pay for things on its own is a problem, not a feature. The question is whether the approval step is something a human can do in one move (read six digits out of an email) or something only a browser can do (a multi-page flow with a bot check).

Not the same as removing the human:

The good version of this keeps the person in the loop and takes them out of the plumbing. A six-digit code proves control of an email address exactly as well as a magic link does, and a human can relay it in five seconds without leaving their editor. The bad version, an agent that can open accounts and spend money with nobody watching, fails a different test entirely.

Test 3: Can you give the agent less than everything?

An agent is a semi-trusted actor, and not because it is careless. It reads untrusted content all day: web pages, issues, emails, other people's documentation. Anything it reads can try to instruct it. That makes a single all-powerful API key a genuinely bad fit, in a way it is not for a human with a password.

So an AI-native product has to let you hand over narrow authority, and the narrowing has to be real: enforced server-side on every interface, and impossible for the holder to widen.

The test: can you give an agent permission to configure things but not to read your data? If the only credential on offer is "an API key", the answer is no, whatever the docs say about best practice.

The sharper follow-up: can a key mint another key with more power than it has? If it can, the scopes are decoration.

Test 4: Is the documentation written for the machine?

A model should not have to infer your product from a hero section. AI-native products describe themselves in a form a program can read: a machine-readable summary of what the product is (llms.txt), tool descriptions that explain when to use a tool and not just its arguments, and instructions aimed at the agent rather than the reader.

That last one is the tell. A product that has thought about agents has written things like "ask the user for their email, never invent one". That guidance exists only because the reader is a program with the ability to guess.

The test: fetch /llms.txt on the vendor's domain. Then read their tool descriptions and ask whether they would help something decide whether to call a tool, or merely how.

Test 5: Are failures legible to a program?

A human reads "Something went wrong" and tries again. An agent needs to know whether to retry, change the request, ask the user for a decision, or give up, and those are four different behaviours. A prose error message with a 500 beside it supports none of them.

The test: cause a failure on purpose (exceed a limit, ask for something your plan does not include) and look at what comes back. A stable code naming the kind of problem (a limit, a permission, a plan, a missing thing) is the difference between an agent that recovers and one that loops.

How to run this on any vendor in ten minutes

No access to their roadmap required:

  1. Fetch https://their-domain/llms.txt. Present and accurate, or missing.
  2. Connect their MCP server, if they have one, and list the tools. Count them against their dashboard navigation.
  3. Ask your agent to do a real task from scratch, not a demo task. Note the first moment it needs you to open a browser. That moment is their honest position.
  4. Try to create the narrowest possible credential. See whether "read configuration, write configuration, never read customer data" is expressible.
  5. Break something deliberately and read the error.

We are deliberately not publishing scores for other form backends here. We would be marking our own competitors' homework, and a snapshot of someone's integration taken today would be unfair within a quarter. The tests are cheap to run; run them.

Where ShipMyForm actually lands

Scored against our own five tests, honestly, including the misses. The detail behind each of these, the full tool list, the scope table and a terminal recording of an agent doing the whole setup, is on our AI-native form backend page.

TestWhere we are
1. Whole product29 MCP tools, covering the account lifecycle, forms, connectors, hosted pages, billing and the inbox. Not complete: see below.
2. Start from nothingsign_up takes an email, verify_account takes the six-digit code you read out. No signup form, no card, no browser.
3. Narrow authorityFour scope presets. wiring can build forms and cannot read a single submission. Enforced on MCP and REST alike; a key can never grant scopes it does not hold.
4. Machine-readablellms.txt, and tool descriptions that say when to use a tool. The server's own instructions tell an agent to ask for the user's email rather than invent one, and to request wiring when it only needs to wire a form up.
5. Legible failuresTool errors carry a code (not_found, plan, limit, scope) so an agent can tell "ask the user to upgrade" from "ask for a wider key" from "stop".

The two it does not pass cleanly

Test 1 has a real gap: OAuth connectors still need a browser. Google Sheets and Notion authorise through a consent screen, and no amount of tool design removes that: the consent screen is the point. An agent can create the form, wire email, Slack, a webhook and a hosted page, and then has to hand you a link for Sheets. We would rather say so than describe the lifecycle as seamless.

Reading submissions is a paid feature. An agent on the free plan can build everything and cannot read what people sent. That is a deliberate commercial decision rather than a technical limit, and it means the free-plan agent experience is "wiring" and not "triage".

There is also a version of test 2 that needs no account at all: point a form at /to/your@email and it works from the first submission. That is the lowest-friction thing an agent can do with us, and for a lot of tasks it is the right one — see a form endpoint with just your email.

Why we think this is worth being deliberate about

Two things changed at once. Agents became competent enough to do multi-step setup work, and people started asking them to. The result is a class of user who will never see your marketing site, judges you entirely on whether they can finish the job, and gives no feedback when they cannot — they just use something else.

The products that read well to that user are not the ones with the most AI inside them. They are the ones whose machine interface is the product rather than a side door, whose permissions can be made small, and whose errors say what went wrong. None of that is a model. It is just taking a new kind of user seriously, which is the only thing "AI-native" should have meant.

If you want the product detail rather than the definition (the tool list, the scopes, what an agent's session actually looks like) it is all on the AI-native form backend page. If you are earlier than that and want to know what the thing being automated even is, start with what a form backend does.

Frequently asked questions

What does AI-native mean?
That an AI agent is treated as a first-class user of the product, not a visitor who has to drive a human interface. In practice: the machine interface covers everything the UI does, an agent can get started without a human in a browser, the authority it is given can be narrowed, the documentation is addressed to machines, and failures come back as codes a program can branch on.
What is the difference between AI-native and AI-powered?
AI-powered describes where the AI sits: inside the product, doing some of the work — classification, summarising, generation. AI-native describes who the product is for: an AI agent is one of its users. A spam filter that uses a model is AI-powered. A product an agent can set up end to end is AI-native. They are independent, and most products claiming the second only have the first.
Is having an MCP server enough to make a product AI-native?
No. An MCP server with three tools that hands the agent a dashboard link for everything else is a machine-readable brochure. The question is coverage: what fraction of what a human can do in the UI can the agent do, and can it get through onboarding, permissions and errors without a person taking over.
How do I test whether a tool is really AI-native?
Ask your agent to accomplish a real task with the vendor, from having no account, and watch for the first moment it has to hand you back a browser. Then try to give it the narrowest possible credential. Where it stops and what it can be trusted with tell you more than any landing page.
Why does an AI agent need scoped permissions rather than an API key?
Because an agent reads untrusted content — web pages, issues, emails — and anything it reads can try to instruct it. That makes a full-access key a liability in a way a human's password is not. Scoping means the agent that builds your forms has no path to your customers' messages, so a bad instruction has nothing valuable to reach.
Can an AI agent use ShipMyForm on the free plan?
For the building work, yes: creating the account, forms, snippets, connectors, hosted pages and stats all work on the free plan with no card. Reading, classifying and exporting submission content over MCP or the REST API is a paid feature, so an agent on the free plan can wire a form up but cannot read what people sent.
What is llms.txt?
A plain-text file at the root of a site that describes what the product is and does, in a form a language model can read without parsing a marketing page. It is the machine-facing counterpart of a sitemap: not a ranking signal, just a way of being described accurately rather than inferred from your own hero copy.

Related guides