
The Machine That Wasn’t One
In 1770, Wolfgang von Kempelen unveiled the “Mechanical Turk”: a chess-playing automaton that stunned European courts. With gears, a cabinet, and an articulated arm that moved its own pieces, it looked like a machine. For decades it beat nearly everyone who sat across from it. Napoleon played it. Ben Franklin played it. It toured the world as “proof” that a machine could think.
Think, it could not. There was a person concealed within the entire time.

These were real chess games: real gambits, real outcomes. The trick was the interface: it made the audience believe the hard part was handled. That someone had solved the intelligence problem underneath.
In 2026, a new generation of AI development tools like Lovable and Replit performs a remarkably similar trick. Describe the application you want, and it appears: working authentication, file uploads, a database, camera permissions, a deploy button. They’ll host it indefinitely, too.
The interface is so smooth that it feels like the hard part has been handled. A student’s homework photo goes in the top, an AI-graded response comes out. It looks like engineering. It looks like infrastructure. It looks like someone solved the problem.
Except… Nobody asks who’s inside the cabinet.
Using Censys Internet-wide scan data, I looked at the population of AI-built web applications that target students and children. Homework helpers, GPA calculators, scholarship matchers, essay graders, admission-chances predictors.
I asked what happens to the data they collect. What I found isn’t malware, phishing, or cybercrime. It’s a story about a competence gate that used to exist and doesn’t anymore, and what rushes through without it.
What Changed, and Why It Matters
Squarespace and Wix made it easy to create a website a decade ago. What they never made easy was building a system that ingests data from a stranger. To do that, you historically needed a form backend, a file storage bucket, a database with a schema, auth with session handling, and somewhere to put credentials. This was a gate to building complex web apps.
Not a legal gate or a safety gate — a competence gate. It filtered out people who might be careless with that data, because they couldn’t get that far.
With an app like Lovable, a person can now type let users upload a photo of their homework and sign in to see results and get working file uploads, an auth database, and camera permissions in simple prompts.
Publishing was never the hard part. The data ingestion was. And now it isn’t.
Censys tags web properties built with these tools using the AI label. That gave me our starting population: 4.37 million websites.
From there, I built fingerprints targeting applications whose own marketing language — in <title> tags and meta descriptions — describes collecting academic data from students: homework submissions, GPA, essays, test scores, extracurriculars, admissions profiles.
None of these appear in any threat catalog. None are flagged as phishing. None serve malware. Nearly all of them are exactly what they claim to be: small, functional, recently deployed education tools. Serving their purpose.
They’re sitting ducks.
More insight on the State of the Internet in 2026
Each year, Censys publishes the definitive view of how the Internet changes. And how we change with it. From AI trends to critical infrastructure and cybercrime. Sign up to get your copy when it publishes this October.
The Numbers
First, I established a definition that qualified live websites for this population: they had to match both an AI-building technology and clear intent to accept student PII (personally identifiable information) or PD (personal data).
Rather than a single monolithic query, the fingerprint is built from three interlocking layers, any of which can be adapted for ongoing monitoring.
75%+ of the population were deployed within the last year.
These aren’t legacy education sites that accumulated over the lifetime of the Internet. They appeared in a wave, in the span of a summer, built with tools that, themselves, exploded into existence.
Intent: What the Site Says It Collects
The HTML <title> tag and <meta> descriptions are the primary text surface. I searched these metadata fields for education-specific intake language: terms like “homework,” “GPA,” “scholarship,” “admission chances,” “essay,” and “K-12,” combined with collection verbs like “upload,” “snap a picture,” or “sign up to see your results.”
Equally important are the exclusions: phrases like “for teachers,” “classroom management,” “lesson planning,” and “LMS” filter out the adjacent population of educator tools that share vocabulary but serve a different audience.
The discriminator is who the site addresses — a tool that says “upload your homework” is speaking to a child; a tool that says “track student progress” is speaking to an adult.
Capability: How the Site Ingests Data
Nearly all of the sites have authentication and account management accepting passwords. Many also prompt the user with surveys that may or may not solicit PII.
Taking it a step further, I looked for direct evidence of data-intake infrastructure: type="file" input elements, accept="image" attributes, getUserMedia calls for camera or microphone access, and references to backend storage services like Supabase or Firebase. I found these in roughly 30% percent of our sample — manifesting as uploads of report cards, handwritten notes, or capture of voice and video recording.
Provenance: the AI Label and What It Reveals
Within our population, Lovable dominates at roughly 229 websites, followed by Replit at 45 — a 5:1 ratio that likely reflects Lovable’s one-click deploy-to-custom-domain workflow, which eliminates the last friction point between prompting an app and putting it on the public Internet. The remaining sites are built on Express, Bubble, Next.js, and traditional stacks where the AI label fired on other signals.
61% of these Lovable-built sites still ship the platform’s default social card — the unmodified og:image, the twitter:site still pointing to @Lovable or @lovable_dev, the stock preview screenshot hosted on Lovable’s R2 bucket.
More than half of the sites in our sample built a system that ingests student data, and never got as far as replacing the default social media card.
What I Found
Each example below was identified through Censys scan data and is presented as it appeared at the time of observation.
hwchecker[.]com

HomeworkChecker explicitly targets K-12 students (down to an AI-generated photo of children) and invites them to upload photos of their homework for instant AI feedback.
Photo upload from a minor is one of the highest-sensitivity intake surfaces on the web. Think about it: a picture of a worksheet can contain the child’s handwriting, full name, school name on the letterhead, and sometimes a face or a bedroom in the background.
The site still loads Lovable’s default OpenGraph image from lovable.dev. Its twitter:site reads @lovable_dev.
Camera access, file upload, AI processing pipeline — the entire website was assembled and deployed without the operator ever customizing the metadata that describes the site to the outside world.
sealedu[.]com

“Sealedu watches your screen, listens to your voice, and guides you like a real professor.”
That’s the meta description, verbatim. The application requests screen-share and microphone access (two of the most sensitive permissions a browser can grant) and positions this as its core feature. A student shares their live screen so the AI can see the problem they’re working on. The AI listens to them talk through it.
In a classroom, this interaction would be governed by FERPA, recorded in an education record, and subject to parental consent.
The site’s og:image is Lovable’s default. The cert was first logged August 13, 2026. No privacy policy is linked in the served HTML. No favicon was even configured — the favicons array is empty, the only site in our sample where the operator didn’t get as far as having an icon. The site asks for a student’s screen and voice and didn’t finish setting up the tab graphic.
collegcalctulater[.]site

The domain is a double misspelling, the kind of typo typically made by someone typing quickly on a phone. The site behind it, “College Compass,” asks students to enter their GPA, essays, extracurricular activities, and work experience to calculate admission chances and match them with scholarships.
Its og:image is still Lovable’s unmodified default preview screenshot. Its twitter:site still reads @Lovable. Again, the operator built a system that ingests student essays and never got as far as replacing the default social card.
joinmasar[.]com

MASAR doesn’t target children directly. It targets universities — and somehow that makes it much worse!
The platform describes itself as “AI-Powered Student Recruitment Simulations” and promises to deliver “qualified, scored leads” by having prospective students play five-minute decision-making scenarios. The AI then profiles each student’s “leadership, critical thinking, risk tolerance, and collaboration skills” and packages the results as a recruitment lead. The student’s name, email, and phone number are, per MASAR’s own copy, “captured naturally” during the simulation.
The product is a lead-generation funnel disguised as a personality quiz, sold to admissions offices at $299–$999/month.
The student thinks they’re exploring whether a university is a good fit. The university is buying a psychometric profile of a minor that the minor didn’t know was being scored.
MASAR also ships Lovable’s default favicon and og:image.
steamgurus[.]org

SteamGurus offers “Personalized STEAM tutoring with AI-powered learning and expert human tutors for K-12 students.” It names the age band in its own title tag, but then leaves it off of the homepage copy.
It addresses students, parents, and educators simultaneously, inviting all three to “be part of the SteamGurus community”. A single intake surface for adults and the children in their care, with no visible differentiation in how data from each group may be handled.
The entire served HTML is 1,430 bytes: a <div id="root"> and nothing else. Every interactive element lives in a JavaScript bundle. The favicon image and hash are Lovable’s default.
The business behind this site may have existed before Lovable – but the AI-generated code is much more recent and already soliciting information from children.
knobee[.]ai and learnwithace[.]study


Knobee invites students to “upload notes, snap pictures of textbooks, or record your questions.” Three intake channels (file upload, camera, microphone) in a single sentence.
LearnWithAce describes itself as “your AI study companion” and advertises “file uploads, camera access, and smart study tools” in the same meta tag. It names the browser permissions it requests the way a feature list names bullet points. Camera access is a selling point, not a disclosure.
Both were built within the same two-week window. Both ship Lovable’s unmodified branding. Both ask students for camera and file access. Neither links a privacy policy.
What Isn’t There — and What Is
There is a conspicuous absence in this data, and it’s worth naming directly: almost none of these sites are malicious in the conventional sense.
I ran the full population against behavioral indicators like ClickFix lures, obfuscated JavaScript, LOLBin command strings, binary delivery, webshell signatures, ephemeral tunnel endpoints, anonymization infrastructure, hidden form fields, insecure submission targets, and certificate hygiene failures. None showed any signal at all. No Fake Captcha. No phishing kits.
Is this… reassuring? It means the problem is outside the vocabulary that security tooling is built to describe.
These aren’t sites that steal data — they’re sites that collect it, earnestly and functionally, built by people who may not know what a privacy policy is for, running on infrastructure they didn’t provision and couldn’t audit. The data goes into a Supabase table or a Firebase bucket that the operator probably accesses through a GUI they’ve never configured beyond defaults.
There is no MITRE ATT&CK technique for “built a homework app in an afternoon and forgot to think about where the photos go.” Yet, the data continues pooling in these systems. And the children don’t experience the distinction between malice and negligence. To them, the exposure is the same.
And then, occasionally, something in the population looks worse than negligent.
lingua-ai[.]aigoconnection[.]com

Most of our examples are hollow — empty scaffolds on shared hosting, with no infrastructure behind them worth scrutinizing. Lingua AI is different. It runs on a dedicated host at 185.106.177.145, and the host is dense.
The application itself is an AI language-learning platform with separate portals for students, teachers, and administrators. The backend is legit: Kestrel (.NET) across five high-numbered ports, an OpenAI API proxy on port 3000, Caddy reverse-proxying traffic on the front. Someone built and operates a multi-service stack that handles student and teacher data across role-segregated access tiers.
The login page also displays, in a yellow tooltip visible to anyone who visits, the default “demo” credentials: admin / 123qwe. The hint helpfully notes that the password can be changed in system settings.
What makes the infrastructure unusual is everything else sharing the host. NPS — a Chinese-origin open-source reverse-proxy tunneling tool that appears frequently in threat research as a post-exploitation utility — runs on ports 8080 and 10080. A DERP relay endpoint (derp1.aigoconnection.com) indicates Tailscale-style VPN tunneling. The parent domain aigoconnection.com also resolves subdomains named remote, m1.dog, m2.dog, and newapi to the same IP.
Over 30 ports are open across non-standard ranges.
The host sits on XNNET LLC, a Hong Kong ASN.
No threat labels or malware indicators are present in the scan data — but the combination of NPS tunneling, a DERP relay, dense multi-port exposure, and an education platform collecting student data on the same host is, at minimum, atypical for a production SaaS deployment.
Make no mistake: The host at 185.106.177[.]145 presents a severely over-exposed attack surface. Student data is entering a system whose boundaries extend well beyond the login page.
The Person Inside the Cabinet
The original Mechanical Turk toured for 84 years before it was destroyed in a fire in 1854. For most of that run, audiences debated whether it was truly intelligent.
Although the “Turk” may have evolved to be newer, shinier, and have a better interface, the structural trick is the same. The gleaming exterior — describe what you want and it appears — makes the audience believe the hard part has been handled.
It makes the audience believe that someone, somewhere, thought about where the homework photos with a name and handwriting are stored.
It makes the audience believe that the dream colleges typed in by a sixteen-year-old are encrypted at rest.
It makes the audience believe that when a child grants camera access to an app that was a prompt forty-eight hours ago, there is an adult on the other end who understands what that permission means.
Only, the cabinet is empty this time.
Kempelen’s “Turk” at least had a chess grandmaster inside. These new machines have no one inside at all. They run exactly as prompted, collecting exactly what they’re told to collect, with no one watching what accumulates.
Censys scans the Internet every day looking for threats. Increasingly, the most consequential things we find aren’t threats in the traditional sense. They’re gaps — places where capability arrived before responsibility, where the tools outran the people using them.
The students don’t know the cabinet is empty. They just see the machine, and they trust it, because it looks like it works.
Appendix: Methodology and CenQL
All queries target Censys web properties and are written in CenQL. The fingerprint is built from three layers, combined with AND.
Layer 1: Content — Education-Specific Language Targeting Students
Searches <title> tags, <meta> descriptions, HTTP body, and hostnames for terms indicating the site targets students and solicits academic data. Exclusions are built into this layer — phrases like “for teachers,” “classroom management,” and “lesson plan” filter out the adjacent population of educator tools that share vocabulary but serve a different audience. Accredited institution domains (.edu, .k12.*.us, .ac.uk) are also excluded.
(
(
(web.endpoints.http.html_tags: "K-12" or web.endpoints.http.html_tags: "K12" or web.endpoints.http.html_tags: "middle school" or web.endpoints.http.html_tags: "high school" or web.endpoints.http.html_tags: "elementary school" or web.endpoints.http.html_tags: "grade level" or web.endpoints.http.html_tags: "schoolwork" or web.endpoints.http.html_tags: "homework")
and (web.endpoints.http.html_tags: "upload" or web.endpoints.http.html_tags: "photo" or web.endpoints.http.html_tags: "snap a picture" or web.endpoints.http.html_tags: "take a photo" or web.endpoints.http.html_tags: "scan your" or web.endpoints.http.html_tags: "camera" or web.endpoints.http.html_tags: "attach")
)
or
(
(web.endpoints.http.html_tags: "admission" or web.endpoints.http.html_tags: "scholarship" or web.endpoints.http.html_tags: "financial aid" or web.endpoints.http.html_tags: "college application" or web.endpoints.http.html_tags: "university application")
and (web.endpoints.http.html_tags: "gpa" or web.endpoints.http.html_tags: "transcript" or web.endpoints.http.html_tags: "extracurricular" or web.endpoints.http.html_tags: "test scores" or web.endpoints.http.html_tags: "sat score" or web.endpoints.http.html_tags: "act score" or web.endpoints.http.html_tags: "admission chances" or web.endpoints.http.html_tags: "your essays")
)
or
(
web.endpoints.http.html_title: "K-12"
or web.endpoints.http.html_title: "homework help"
or web.endpoints.http.html_title: "homework helper"
or web.endpoints.http.html_title: "homework solver"
or web.endpoints.http.html_title: "admission chances"
or web.endpoints.http.html_title: "chance me"
or web.endpoints.http.html_title: "scholarship matcher"
or web.endpoints.http.html_title: "gpa calculator"
or web.endpoints.http.html_title: "essay grader"
or web.endpoints.http.html_title: "ai tutor"
or web.endpoints.http.html_title: "study buddy"
)
or
(
(web.endpoints.http.html_tags: "essay" or web.endpoints.http.html_tags: "assignment" or web.endpoints.http.html_tags: "coursework" or web.endpoints.http.html_tags: "book report" or web.endpoints.http.html_tags: "lab report")
and (web.endpoints.http.html_tags: "grader" or web.endpoints.http.html_tags: "solver" or web.endpoints.http.html_tags: "humanize" or web.endpoints.http.html_tags: "undetectable" or web.endpoints.http.html_tags: "paraphras" or web.endpoints.http.html_tags: "rewrite" or web.endpoints.http.html_tags: "instant feedback")
and (web.endpoints.http.html_tags: "student" or web.endpoints.http.html_tags: "school" or web.endpoints.http.html_tags: "homework" or web.endpoints.http.html_tags: "K-12" or web.endpoints.http.html_tags: "high school" or web.endpoints.http.html_tags: "teacher")
)
or
(
(web.hostname: "homework" or web.hostname: "hwcheck" or web.hostname: "studybuddy" or web.hostname: "gradecalc" or web.hostname: "collegecalc" or web.hostname: "chanceme" or web.hostname: "scholarship" or web.hostname: "essaygrader" or web.hostname: "calctulater" or web.hostname: "scholorship" or web.hostname: "homwork")
and (web.endpoints.http.html_tags: "student" or web.endpoints.http.html_tags: "school" or web.endpoints.http.html_tags: "homework" or web.endpoints.http.html_tags: "college" or web.endpoints.http.html_tags: "scholarship" or web.endpoints.http.html_tags: "tutor")
)
)
and not web.endpoints.http.html_tags: "for teachers"
and not web.endpoints.http.html_tags: "for educators"
and not web.endpoints.http.html_tags: "for administrators"
and not web.endpoints.http.html_tags: "classroom management"
and not web.endpoints.http.html_tags: "lesson plan"
and not web.endpoints.http.html_tags: "curriculum builder"
and not web.endpoints.http.html_tags: "professional development"
and not web.endpoints.http.html_tags: "teacher dashboard"
and not web.endpoints.http.html_tags: "school management"
and not web.endpoints.http.html_tags: "school administration"
and not web.endpoints.http.html_tags: "LMS"
and not web.endpoints.http.body: "for teachers"
and not web.endpoints.http.body: "for educators"
and not web.endpoints.http.body: "lesson planning"
and not web.endpoints.http.body: "classroom management"
and not web.endpoints.http.body: "teacher dashboard"
and not web.endpoints.http.body: "manage your classroom"
and not web.endpoints.http.body: "track student progress"
and not web.endpoints.http.body: "teacher resources"
and not web.endpoints.http.body: "professional development"
and not web.endpoints.http.html_tags: "enterprise-grade"
and not web.endpoints.http.html_tags: "professional-grade"
and not web.endpoints.http.html_tags: "pro-grade"
and not web.endpoints.http.html_tags: "military-grade"
and not web.endpoints.http.html_tags: "investment-grade"
and not web.hostname =~ "\\.edu$"
and not web.hostname =~ "\\.k12\\.[a-z][a-z]\\.us$"
and not web.hostname =~ "\\.ac\\.uk$"
Layer 2: Provenance — AI-Built Applications
Matches web properties tagged by Censys as built with AI development tools. Within our population, Lovable accounts for roughly 5× more properties than Replit.
and (
web.labels.value = "AI"
or host.services.labels.value = "AI"
)
Layer 3: Data Input Methods — File Upload, Camera, and Storage
Isolates the subset showing direct evidence of file-intake or device-access capability in served HTML or metadata.
and (
web.endpoints.http.body: "type=\"file\""
or web.endpoints.http.body: "accept=\"image"
or web.endpoints.http.body: "getUserMedia"
or web.endpoints.http.body: "capture=\"camera\""
or web.endpoints.http.body: "capture=\"environment\""
or web.endpoints.http.body: "navigator.mediaDevices"
or web.endpoints.http.body: "getDisplayMedia"
or web.endpoints.http.body: "supabase"
or web.endpoints.http.body: "firebasestorage"
or web.endpoints.http.body: "storage/v1/object"
or web.endpoints.http.body: "dropzone"
or web.endpoints.http.body: "FileReader"
or web.endpoints.http.body: "readAsDataURL"
or web.endpoints.http.body: "createObjectURL"
or web.endpoints.http.html_tags: "upload"
or web.endpoints.http.html_tags: "camera access"
or web.endpoints.http.html_tags: "snap a picture"
or web.endpoints.http.html_tags: "take a photo"
or web.endpoints.http.html_tags: "upload a photo"
or web.endpoints.http.html_tags: "record your questions"
or web.endpoints.http.html_tags: "file upload"
or web.endpoints.http.html_tags: "drag and drop"
)
Supplementary: default-branding detection
Appended to identify Lovable-built sites that still ship unmodified platform defaults
and (
web.endpoints.http.body: "twitter:site\" content=\"@Lovable"
or web.endpoints.http.body: "twitter:site\" content=\"@lovable_dev"
or web.endpoints.http.body: "opengraph-image-p98pqg"
or web.endpoints.http.body: "id-preview-"
or web.endpoints.http.body: "pub-bb2e103a32db4e198524a2e9ed8f35b4.r2.dev"
)

