InternetSocial MediaTechnology

Social Media Child Safety Features Fail: Shocking Study Finds 60% Broken in 2026

Social media child safety features mostly fail, a new study finds. See which platforms failed worst and what parents should do now.

For years, the big platforms have told parents the same story: don’t worry, we’ve built tools to keep your kids safe. Search filters. Messaging blocks. Time limits. Anti-bullying prompts. It’s a long list, and it sounds reassuring on paper.

A new study says most of it doesn’t actually work.

Researchers at the Cybersafety Research Center, a joint effort between New York University and Northeastern University, spent six months testing 86 social media child safety features across Instagram, TikTok, Snapchat, and YouTube. They didn’t just read the marketing copy. They built fake teen and adult accounts, used the platforms the way real kids and real predators actually use them, and tracked whether each tool did what the company claimed.

The result, published in a report called “Broken, Buried, or Missing,” is hard to spin. Of the 86 features tested, 51 failed. That’s roughly 60%. Only 35 worked as advertised and were actually visible or usable by a young person on the platform.

This isn’t a small methodological quibble. Some of the failures are the kind that make you want to hand your kid a flip phone. A TikTok account registered to a minor searched for content about self-harm and eating disorders, and instead of blocking the search, the app started suggesting related terms pulled straight from pro-anorexia communities. That’s not a safety feature falling short. That’s the opposite of one.

Below is a full breakdown of what the researchers found, which platforms performed worst, why these tools keep failing, and what parents can realistically do about it.

What the Study Actually Tested

The Cybersafety Research Center didn’t invent a new theory of online harm. They built on existing frameworks researchers have used for over a decade to categorize the risks kids face on the internet, then added two categories of their own to reflect how modern apps actually work.

Their testing covered five broad categories:

  • Content risk — exposure to harmful material like self-harm content, graphic violence, or pro-eating-disorder posts
  • Contact risk — unwanted contact from strangers, including adults messaging minors
  • Conduct risk — bullying, harassment, and harmful behavior between users, including anti-bullying prompts
  • Circulation risk — a new category the researchers added, covering what happens when a minor’s content spreads beyond its intended audience (their gymnastics-team photo example, where an unrelated adult stumbles onto a teen’s posts, is a good illustration)
  • Compulsivity risk — also newly added, covering design patterns like autoplay and infinite scroll that make it hard for anyone, let alone a teenager, to log off

To test all of this, researchers set up dummy accounts posing as children of different ages, alongside adult accounts, and ran through three separate scenarios for each child safety feature:

  1. A child using the platform normally, encountering the feature the way it was designed to appear
  2. A teen actively trying to get around the safety feature, the way real teenagers do
  3. An adult attempting to exploit gaps in the system, such as messaging or searching for a minor’s account

A feature only counted as a “success” if it functioned as the company claimed and it was actually shown to, or usable by, a child account during testing. If a tool existed but was buried three menus deep in settings, or technically worked but never surfaced when it should have, it didn’t pass.

Instagram had the most features under review, with 29 tested. TikTok had 24, YouTube had 22, and Snapchat had fewer still. That range alone tells you something: platforms aren’t building safety infrastructure at the same pace, even though they’re all marketing themselves as safe spaces for teens.

The Results, Platform by Platform

The headline number is that 60% of all tested features failed. But the failure rate wasn’t evenly distributed. Some platforms did meaningfully worse than others.

Snapchat: 73% Failure Rate

Snapchat came out worst in the study. Researchers found that adult test accounts were able to search for, locate, and message a minor’s account with essentially no friction. No warning prompt, no restriction, nothing standing between a stranger and a child’s inbox. For an app built around private, disappearing messages, that’s about as bad a result as you can get on contact risk.

Instagram: 66% Failure Rate

Instagram, Meta’s flagship app for teens, didn’t fare much better. One of the more telling failures involved the platform’s “pause to rethink” prompt, a feature meant to nudge users to reconsider before posting something unkind. In testing, comments containing bullying language and insults between two teen accounts didn’t trigger the prompt at all. Meta’s explanation was that the feature isn’t designed to appear when the commenter and the poster already follow each other, which, if you think about it for a second, describes the vast majority of bullying that actually happens between kids who know each other.

Instagram’s autocomplete search also proved easy to trick. When testers began typing a search related to eating disorders, the platform’s own autocomplete suggested deliberately misspelled variations, the exact kind of spelling tricks that pro-eating-disorder communities use to dodge keyword blocklists.

YouTube: 55% Failure Rate

YouTube landed in the middle of the pack. Its failures skewed toward features that existed but weren’t consistently surfaced to the accounts that needed them, rather than tools that were completely absent. Still, more than half of what was tested didn’t hold up.

TikTok: 50% Failure Rate

TikTok had the best relative performance of the four platforms, but “best” here still means half of its safety features failed. And one of its failures was arguably the most disturbing finding in the entire report. A minor test account that searched for material related to self-harm and disordered eating wasn’t blocked. Instead, TikTok’s search function began actively recommending related terms, including phrases referencing self-cutting and language used by pro-anorexia communities to talk about hiding food from parents. Researchers said the harmful suggestions surfaced almost immediately, within a single search.

Why Search Filters Keep Failing

A recurring theme across every platform was how easy it was to slip past keyword-based content filters. Snapchat’s filtering system is a good example: researchers found that one misspelled version of a harmful search term got blocked, while the correctly spelled version of the same term returned results with no restriction at all. In multiple cases, researchers said it took less than three minutes to find a working bypass for a filter that was supposedly built to stop exactly that kind of search.

That’s the core problem with relying on keyword lists to catch harmful content. Teenagers, and honestly anyone who’s spent time in online communities built around evading moderation, learn the workarounds fast. A filter that only catches the “correct” spelling of a dangerous term isn’t a safety feature so much as a speed bump.

The Categories That Failed Hardest

Not every type of safety feature failed at the same rate. Some categories were consistently weak across all four platforms, which suggests the problem isn’t isolated to one company’s engineering team, it’s closer to an industry-wide pattern.

Conduct and anti-bullying tools performed the worst of any category. Every single conduct safeguard designed to prevent cyberbullying failed on at least one platform, and in several cases failed across all four. These are the tools meant to catch harassment, insults, and pile-ons between users, and they were among the least reliable features tested anywhere in the report.

Beyond conduct tools, researchers grouped the failures into three buckets:

  • Broken features — the tool existed and was reachable, but simply didn’t do what it claimed
  • Buried features — the tool worked in theory, but was hidden deep enough in settings menus that a realistic teen user would never find it
  • Missing features — nine features fell into this category entirely; researchers could never get them to trigger at all, no matter how they tried

That last group is worth sitting with. Nine tools that platforms had publicly promoted as protecting kids simply never activated, under any of the three testing scenarios researchers used.

What Actually Worked

It’s not all bad news, and it’s worth being fair about what the research found on the other side of the ledger. The features that consistently succeeded shared two traits: they were turned on by default, and they removed risky functionality entirely for the youngest users rather than trying to filter or moderate it after the fact.

In other words, the safety features that worked weren’t the clever ones. They weren’t the AI-powered content classifiers or the nuanced keyword systems. They were the blunt tools: turning off public search for accounts under a certain age, disabling messaging from non-followers by default, or removing features like autoplay entirely for young teen accounts rather than trying to limit how long they could use it.

That’s a meaningful finding for parents and for regulators, because it points to a fairly simple design principle: safety that depends on a child finding and activating a setting is safety that mostly doesn’t happen. Safety that’s built into the default experience, especially for accounts registered as under 16, held up far better under testing.

How the Platforms Responded

None of the four companies fully accepted the findings. Spokespeople for Meta, Snap, and YouTube pushed back, arguing in statements that their features function as intended or that the researchers’ testing didn’t reflect how real teens actually use the apps. A YouTube representative pointed to internal survey data claiming that a large majority of parents using its supervised account tools said the tools gave them more confidence in their child’s online safety.

That pushback is worth taking seriously, and it’s also worth noting that at least one outside newsroom, The New York Times, ran its own tests based on the study’s methodology and reported it was able to replicate several of the findings, including Snapchat’s messaging gap between adult and minor accounts.

It’s also worth putting this study in context. Both Meta and YouTube’s parent company were found liable earlier this year in separate legal proceedings related to designing products that intentionally fostered addictive use among young people. All four companies named in the report are currently facing thousands of lawsuits from families, school districts, and state attorneys general alleging that their platforms caused harm to minors. The companies deny the allegations in those cases.

Why This Study Matters Beyond the Headlines

This report lands at a moment when governments are moving faster than platforms on youth safety. Australia’s under-16 social media ban, which took effect late last year, requires companies to take “reasonable steps” to verify user age, and regulators there have already doubled maximum fines after finding that plenty of underage users were still slipping through. The UK has moved in a similar direction with its own restrictions on under-16 access to major platforms.

Age verification and child safety features are related but separate problems. A platform can get age checks right and still fail at protecting the kids who are legitimately using the app at an approved age, which is exactly the gap this study is pointing at. You can be 15, using Instagram exactly the way you’re supposed to, with an account that was verified as belonging to a 15-year-old, and still end up in a comment thread with bullying language that never triggers a single intervention.

That distinction matters for how parents think about risk. Verifying that your child is using an age-appropriate account doesn’t mean the account is actually safe. The safety has to be built into the product itself, not just gated at the door.

Practical Steps for Parents Right Now

Waiting for platforms to fix these gaps on their own isn’t a great plan, especially given how many of the failures in this study were things companies had already promoted as working. Here’s what actually moves the needle based on what the researchers found held up under testing:

  1. Use the most restrictive default settings available, especially for accounts under 16. The study found that removing risky functionality outright, rather than relying on filters, was consistently the most effective approach.
  2. Turn off public search and discoverability on your child’s account rather than trusting a messaging filter to catch unwanted contact after the fact.
  3. Don’t rely on keyword-based content filters as your main line of defense. The research found these were bypassed in minutes, often with something as simple as a deliberate misspelling.
  4. Check in on conduct, not just content. Anti-bullying and pause-to-reflect prompts were the weakest category across every platform tested, so don’t assume the app is catching harassment between kids who already follow each other.
  5. Talk to your kids about what “supervised” accounts do and don’t catch. A feature being turned on doesn’t mean it’s functioning, and teens who understand the gaps are better positioned to flag problems themselves.
  6. Revisit settings periodically. Platforms update features often, and a setting that worked six months ago may have shifted, been renamed, or been quietly deprecated.

None of this replaces the need for platforms to build safer defaults from the ground up. But until that happens at scale, treating every advertised safety feature as fully reliable isn’t a safe assumption.

The Bigger Picture

What makes this research land differently than the usual round of tech criticism is the specificity. This wasn’t a survey asking parents how safe they felt their kids were. It was a hands-on audit, feature by feature, testing whether the tools platforms spent years promoting actually functioned the way the companies said they did. Sixty percent of them didn’t.

The Cybersafety Research Center’s findings echo a broader pattern researchers have documented for years: platform-based safety tools are often built without enough public evaluation or input from the young people they’re meant to protect, which tends to produce systems that look good in a press release and fall apart under real-world use. That’s precisely what this study set out to measure, and precisely what it found.

For families navigating this without waiting on regulation or corporate goodwill, the honest takeaway is that social media safety for children currently depends more on aggressive default settings and ongoing parental attention than on trusting the safety features these platforms advertise. Read the full CNN report on the findings and Northeastern University’s coverage of the research for more detail on the methodology and the platforms’ responses.

Conclusion

The Cybersafety Research Center’s audit of 86 social media child safety features across Instagram, TikTok, Snapchat, and YouTube found that 51 of them, roughly 60%, failed to work as the platforms advertised, with Snapchat performing worst at a 73% failure rate and TikTok performing best at 50%. Anti-bullying and conduct tools failed most consistently across every platform, while keyword-based content filters proved easy to bypass, sometimes in under three minutes. The features that did succeed shared a common trait: they removed risky functionality entirely by default rather than relying on filters or settings a child had to find and turn on themselves. With companies disputing the findings even as they face mounting lawsuits and new government regulation around youth online safety, parents are left with a clear, if uncomfortable, conclusion: advertised safety features are not a substitute for restrictive defaults and active oversight.

5/5 - (3 votes)

You May Also Like

Back to top button