All resources
Spam protection

Why Your Contact Form Suddenly Started Getting Spam

Nothing about your site changed — a crawler found your form. How bots discover forms, why it starts all at once, what the four kinds of junk are actually after, and which defence works for each.

The ShipMyForm team

· 7 min read

Nothing about your site changed. A crawler found your form. Form-spam bots sweep the web for <form> tags and recognisable input names, submit to everything they find, and keep whatever responds. The volume arrives all at once because discovery is the event, not anything you did, and it stays roughly steady afterwards because your endpoint is now on a list that gets traded.

Full disclosure: ShipMyForm is our product, and it does the filtering described near the end. The mechanics in the first half are true of any form on any backend.

Nobody chose you

The most useful thing to understand is that this is not aimed at you. There is no person reading your site and deciding to target it. There is a crawler, and the crawler is indiscriminate.

It works roughly like a search engine with worse manners. It follows links, reads sitemaps, and parses each page it reaches for a <form> element. If it finds one, it reads the input names, and name, email, message, subject and phone are a near-universal vocabulary. That is enough to fill the form convincingly. It submits, records whether the submission appeared to succeed, and moves on.

Everything that follows comes from that one fact. The junk is generic because the sender never looked at your site. It is grammatically odd because it was templated. It keeps coming because you are a row in a database, and the database gets sold.

Why it started on an ordinary Tuesday

People usually assume they broke something, or that a particular person found them. Neither is typical. The common triggers are all moments of discovery:

  • The page got indexed. A new page or a new form on an old page is found soon after it becomes reachable.
  • Something linked to it. A directory, a forum answer, a repository, a client's site, or a sitemap update.
  • You reused a template. A theme or starter used across many sites gets found once, and every site using it inherits the attention.
  • The address behind the form leaked separately. Different channel, same day, easy to confuse.

None of these change anything about the form. They change who knows it exists, which is why the onset feels so abrupt — one day nothing, the next day a steady trickle that never stops.

What the four kinds are actually after

This is the part that matters, because the right defence depends entirely on what the sender wants. Almost everything you receive is one of four things.

Link dropping. A message carrying URLs, usually several, hoping the text gets published somewhere a crawler will read: a comment, a testimonial wall, a ticket system with public pages, or an email that ends up in a searchable archive. The giveaway is that the message is addressed to no one and makes no request.

Probing. Very short, often nonsense. One word, a stray character, a test address. Nobody wants anything from you yet — the bot is establishing whether the endpoint accepts submissions at all and whether anything comes back. A form that answers probes politely gets promoted to a better list.

Reply bait. A plausible, mildly flattering business enquiry with no specifics. The payload is not the message; it is your reply. A thread with a real human is the opening move for invoice fraud and business-email compromise, and it is also how a list seller proves an address is read by a person rather than a filter. This is the only category where responding makes things materially worse.

Outsourced outreach. Not automated at all. A person paid by volume, offering SEO, redesign or development services, working from a scraped list of forms. Filters struggle with it precisely because it is a human typing into your form the way a customer would.

Why the categories are the point:

A honeypot catches the first two outright and costs real visitors nothing. Reply bait often passes every mechanical check, because it is written to look like a customer. Outsourced outreach passes all of them, because it is a person. If you only measure "how much got blocked" you will think you have solved a problem that has four separate shapes.

Why a CAPTCHA is the wrong first move

The instinct is to put a challenge in front of the form. It does work, and the bill goes to the wrong people: every real visitor pays in time and irritation so that a minority of automated traffic is inconvenienced, and some of your actual customers abandon the form rather than solve a puzzle.

It is also the layer bots have adapted to most. Solving services exist and are cheap. A challenge is worth keeping in reserve for a form that is genuinely under sustained attack; it is a poor first response to ordinary background noise. Stopping form spam without reCAPTCHA covers the cheaper layers in order, and Turnstile is the least intrusive challenge if you decide you need one.

What actually stops each kind

In rough order of cost to the visitor, which is the order worth trying them in. None of this requires the visitor to do anything.

A honeypot field. A hidden input a human never sees and never fills. Anything that fills it is automated, with no ambiguity at all — which is why it is usually treated as conclusive rather than as a point on a scale. Free, invisible, and it removes the bulk of naive traffic. Every form on this site has one.

A time trap. Humans take time to write. A submission that arrives a second or two after the page rendered was not typed by a person. Cheap, invisible, and it catches bots that render the page properly enough to skip hidden fields.

Content scoring. Count the links. Look for the vocabulary that only ever appears in spam. Check whether the address is a throwaway domain. Each signal is weak alone and the combination is not: several weak signals together are a reliable verdict, and the ones that matter here are link density and disposable senders.

Rate limiting. One source submitting repeatedly is a different event from many sources submitting once. This is what separates an attack from background noise, and it is the layer that matters most when a form is actually targeted.

A blocklist. Once you have seen a pattern — a domain, a phrase, a country you do not serve — you should be able to say so permanently rather than relying on a score.

The counterintuitive part: never tell them it failed

A rejected bot learns something. An error message, a visible challenge appearing on retry, or a submission that obviously vanishes are all feedback, and feedback is what lets an operator adapt the template until it gets through.

So the right behaviour on detection is to accept the submission, return exactly the success the sender expected, and quietly file it away from your inbox. The bot records a win and moves on; nobody tunes anything. This is why a good filter looks, from the outside, like it is not there — and why "I still see them in a spam folder" is a feature rather than a failure. You want the record, without the interruption.

Telling background noise from an attack

Most forms settle into a steady low trickle. Worth escalating when:

  • volume jumps by an order of magnitude rather than drifting up,
  • submissions share a source or arrive in bursts seconds apart,
  • the content mutates between attempts, which means someone is iterating against your filter deliberately,
  • or real submissions start getting lost in the volume.

The first three are when a visible challenge earns its place. The last is the only one that is genuinely urgent, because the cost has stopped being annoyance and started being missed customers.

A ten-minute triage

  1. Add a honeypot if the form does not have one. Biggest single reduction for zero visitor cost. The HTML form checker flags a form that is missing one.
  2. Stop replying to anything unsolicited, especially the polite ones. That is the category where a reply is the objective.
  3. Read what you are receiving and sort it into the four kinds above. It tells you whether you have a bot problem or a people problem, and they have different answers.
  4. Turn on rate limiting before you turn on a challenge.
  5. Only then consider a challenge, and only on the form that is actually being hit.

Where this sits with us

On ShipMyForm this is the default rather than a setting. A tripped honeypot is conclusive and the submission is never stored at all. Short of that, a submission completed implausibly fast, carrying several links, or coming from a throwaway domain accumulates score, and the total decides whether it reaches your inbox or lands in a spam folder you can review. Either way the sender sees an ordinary success, for the reason above.

The part worth knowing if you are being hit hard: neither spam nor borderline submissions count against your monthly quota. That is deliberate rather than generous. When a form is under attack the quota is the thing being attacked, so counting the attack against you would hand the attacker the win.

Two things we deliberately do not do: make you configure any of it before the form works, and put a puzzle in front of your visitors by default. How ShipMyForm fights form spam has the layers in full, and what a form backend is covers where the filtering sits in the path a submission takes.

Frequently asked questions

Why am I suddenly getting spam on my contact form?
Because a crawler found it, usually within days of the page being indexed or linked from anywhere public. Discovery is the event, not anything you changed. Form-spam bots crawl the web looking for <form> tags and recognisable input names, submit to everything they find, and keep the addresses that respond. Once your endpoint is on one of those lists it gets traded, which is why the volume arrives all at once and then stays roughly steady.
How do bots find my contact form if I never shared the URL?
They do not need you to share it. A crawler follows links, reads sitemaps, and parses any page it reaches for a <form> element and input names like name, email and message. Nothing about the form has to be advertised. Pages reachable from a sitemap, a backlink, a public repository or a site template reused across many domains all get found without anyone linking to them deliberately.
Will a CAPTCHA stop contact form spam?
It will stop some of it, and it charges every real visitor for the privilege. A visible challenge is the most expensive defence you have, paid by the 99% of people who are not bots, and the cheaper layers — a honeypot field, a time trap, content scoring and rate limiting — catch most automated traffic before a challenge is needed. Keep a challenge in reserve for a form that is genuinely under attack rather than making it the first move.
Why do I get spam even though my form has no email address on the page?
Because the spam is going through the form, not to a harvested address. The bot posts to the form endpoint and your backend delivers the message to you. Hiding your address helps against address scrapers and does nothing against form submitters, which is why a form can be the only spam channel into an otherwise quiet inbox.
Should I reply to a spam form submission?
No. For one category of form spam the reply is the whole point: the message is bait designed to start a thread with a real person, which is where invoice fraud and business-email compromise begin. A reply also confirms your address is read by a human, which is worth more to a list seller than the submission itself.
Is a honeypot field enough on its own?
It catches the naive majority and costs real visitors nothing, which makes it the best first layer. It will not stop a bot that renders the page and skips hidden fields, so it is a layer rather than a solution. Pair it with a time trap, content scoring and rate limiting, and the combination handles almost everything without a visible challenge.

Related guides