Hacker Newsnew | past | comments | ask | show | jobs | submit | rebeccaskinner's commentslogin

on-device filtering is the right way to go, but it shouldn’t be a matter of sending metadata about the user to a service- instead services should provide standard metadata about the content to the device and the device can choose what, if anything, to display.

This preserves privacy better by keeping more information about the user local, and gives people better tools to decide what metadata categories they want to filter- for their kids and for themselves.


That's much harder to implement. If you ban advertising eg. online casinos to underge users, the service would have to send two different ads (one 18+ and one for younger people too), and the browser would have to decide which one to play. Same for eg searching on google, browsers would have to filter out every <div> with adult content, meaning half the page would be empty, but with a browser "underage" flag, google could just turn on safe search and not allow turning it off.


Plus, as soon as the user clicks on one of these optionally filtered elements, the service will know how exactly what browser is filtering.


That solution doesn't actually work, and it angers me that people can't see why it doesn't work. That solution requires that the entire website must be child-safe or none of it. That solution requires that Tumblr must ban porn if Tumblr has any underage users. The alternative would be that Tumblr would randomly not load for underage users because some recommendation or ad would be over 18, and so it would rapidly have none.


Why must an age rating apply to an entire website?

As one example, the Internet Content Rating Association (ICRA) (and to a lesser extent the prior Recreational Software Advisory Council (RSACi)) had a rating scheme that allowed sites to provide a default rating label that was then overridden for individual pages and resources using <meta> tags and RDF-based labels. For example, Tumblr could set a family-friendly default rating and append a different rating on specific user pages, images, or ads.

An even more granular scheme could be applied via HTML attributes which would allow an individual section, image, link, or text snippet to be marked e.g. as "sexual", "violent", "substance use", "spoiler", or even "unknown" for unreviewed user-provided content. Then leave it up to web browsers to choose whether and how to render elements with these attributes. (Hidden entirely? With censor bars? Pixelated?)

There would be substantial logistical and regulatory challenges to get websites to comply, but it doesn't seem substantially harder than the current age verification schemes.


Because let's say you have a website like Reddit with a mixture of 18+ and 13+ stuff. Something that's 18+ but not explicit gets really popular. Pornhub will now livestream congressional debates, someone posts a link to this, and it gets fifty zillion upvotes as people can't resist commenting "wtf". Now you have a dilemma: do you show it on the front page or not? If you show this on the front page and it's 18+, it'll lock children out of the front page, which for many of them means they're locked out of the entire site because that's the way they know to access the site. But if you don't, then the law is forcing you to censor your front page for everyone, even for adults.

The solution is obvious: you show it on the front page if the user is 18+. However, your proposal deliberately forbids this and says the front page must be the same for everyone. Which leaves the other two bad options.


Haven't paid video streaming services solved that one already? Admittedly not with age verification, but you can set up child accounts. I presume Disney+ is ok here as they want to be in the family-friendly market but also offer content meant for young adults.


That's exactly what the law will do, but on the device. The parents will input the child's age during device setup, and then every service will query the age range based on that and work in "child account" mode.

Without any requirement for verification, it'll be completely up to the parents to decide what their child will see, while the services will only get the minimum information needed. It'll basically make parental control easy, but still in the parent's control.


> If you show this on the front page and it's 18+, it'll lock children out of the front page

I don't understand why you think this is true when I specifically described a label that is applied to "an individual section" of the page.

To be more explicit: Today, each post on the Reddit front page appears in its own container element. If an 18+ post's container element could have some kind of "adult-content" attribute set then some web browsers could render the whole front page normally with the exception of that individual element. Perhaps they could black it out, collapse it, display it with a pixelated overlay and a "Request Access" button, whatever the browser supports and the device administrator prefers.

> The solution is obvious: you show it on the front page if the user is 18+. However, your proposal deliberately forbids this and says the front page must be the same for everyone.

Why do you think that my proposal requires that the front page must be the same for everyone? Providing different views to different users is the entire point of every child safety law and rating scheme I've ever seen. I'm merely saying that my preferred approach would be to mandate that potentially objectionable content is semantically annotated in the HTML somehow, so that each user's browser can apply whatever view restrictions (if any) the device's owner wishes. I think this is better than mandating that web browsers send personal information to every web site they visit and also mandating that sites pre-filter the content they send back.


Won't that effectively leak the user's registered age bracket anyway? And if so, what advantage does it has over the current law?


Good questions!

First, leaking a device's content filtering settings is not the same as leaking a user's age bracket. For example, if the server identifies a web browser that isn't loading tags annotated with "sexual-content" that might indicate that the user is a child but could equally well indicate that they're a corporate office worker, religious, or anyone else who prefers not to see such content at the moment.

Second, if the standard defined multiple semantic tags or levels (e.g. violence, extreme violence, substance use, unknown, etc) then this gives more granular control to device administrators who may care more about some types of content than others.

Third, web browsers wouldn't even have to leak the content filtering settings. For example, perhaps an administrator could configure the web browser to render content overlaid with a semi-opaque blur, in which case the server would not know.

Fourth, and this is more philosophical, asking websites to provide more information to the device so that an admin can do their own filtering leaves the choices to them. Whereas if the only way a device admin can filter content is to submit the age of the device's user to a third party so the third party can decide how to pre-filter the content it sends back, that gives an uncomfortable amount of control to platforms and regulators.


Jellyfin will let you set up alternate metadata providers and select between different episode ordering if available (https://forum.jellyfin.org/t-resolved-dvd-order-instead-of-a...). For tv shows especially TVDB will tend to have alternate orderings most often, and TMDB will occasionally have alternate orderings available.


I don’t think pricing is the only problem with theaters. Especially over the last 15 years or so they’ve been increasingly competing with alternative ways of watching movies, and for a lot of people watching at home wins at any price.

Fancier theaters like the Alamo draft house seem to be trying to complete with watching at home in some ways, but for the most part theaters seem to just be doubling down on the parts if the experience that were already decisive- namely getting louder and adding bigger screens. That might tempt the people who already like what theaters have to offer into going slightly more often, but I think it’s made even more people stop going at all.


It is pretty riddled with ads though. By default the Home Screen shows the tv all and it’s completely crammed with ads.


I sometimes talk with ChatGPT in a conversational style when thinking critically about media. In general I find the conversational style a useful format for my own exploration of media, and it can be particularly useful for quickly referencing work by particular directors for example.

Normally it does fairly well but the guardrails sometimes kick even with fairly popular mainstream media- for example I’ve recently been watching Shameless and a few of the plot lines caused the model to generate output that hit the content moderation layer, even when the discussion was focused on critical analysis.


Interesting. Specific examples of what was censored?


I was a pretty active member in the comments for a long time and left a few years ago after getting chastised by a moderator and accused of spamming for sharing a link to a blog post I had written, even though the content was purely technical, not promoting any product, and does not contain ads or monetize content in any way.

My impression is that the site was actively looking for any possible reason to remove people from the platform. It’s their site to moderate as they wish, but that’s not a community I want to continue participating in.


You did not share a link to a blog post. The title was "Effective Haskell is a hands-on practical book way to learn Haskell. No math or formal CS needed" and it linked to the site advertising your book for sale. I removed it because we don't get good discussions out of ads.


I shared the story as I remember it. Memory is imperfect. It's been years since I deleted my account, and I don't have the luxury of access to server or moderation logs.

What I do remember unambiguously is being an active member of the site, contributing regularly and in good faith, being accused of spamming, and the general feeling of hostility that I got from the site.


You got a DM and email with the title and URL when your story was removed. This would've been 2023-08-03 with the subject "Your story has been edited by a moderator", if you want to look back: https://github.com/lobsters/lobsters/blob/86e1d0b6ac6bac5210...

But you're correct on the second part, there isn't a level of activity that entitles anyone to post a sales page with nothing to discuss on it. Your activity was taken into account, though. Typically if a new user's first activity is to post an ad I'll also ban the site or user. I understand the rules aren't as permissive as you wanted, but ads don't start good discussions.


Your post title doesn't sound like spam to me. Moreover, the link you originally shared https://pragprog.com/titles/rshaskell/effective-haskell/ looks informative enough for discussion; it links to PDFs of the actual content to read. HN, for example, didn't delete it when you posted it https://news.ycombinator.com/item?id=36987260

IMO, the lobste.rs admin's assertion that the post had "nothing to discuss" is a misjudgment that undercuts the rest of their rationalization. My guess is that they're looking for a win on technicality, instead of addressing the myriad of concerns raised elsewhere in this thread.



That website was actually https://web.archive.org/web/20230804152033/https://effective... back in 2023. Even less sale-sy than the HN link.

I don't think the technical win you want is possible or even worth it.


I don't know why you think I want a "technical win" from you, but I'm not seeking your approval. I corrected your mistake about the URL and the policy, like I corrected the author's mistake about what I removed. If you and other sites prefer different policies, it's no skin off my nose.


The vanity of internet moderators never cease to amaze me.


The only acceptable number of ads is zero.


Good luck with that, no company anywhere offers no ads.


Which is why things like The Pirate Bay remain popular.


I’m finishing up Haskell Brain Teasers (https://pragprog.com/titles/haskellbt/haskell-brain-teasers/)

It’s much shorter than my first book, Effective Haskell, and leans more advanced, especially toward the end. Although the format is puzzle focused I’m trying to avoid simple gotcha questions and instead use each puzzle as a launchpad for discussing how to reason about programs, design tradeoffs, and nuances around maintainability.


I worked on a foveated video streaming system for 3D video back in 2008, and we used eye tracking and extrapolated a pretty simple motion vector for eyes and ignored saccades entirely. It worked well, you really don't notice the lower detail in the periphery and with a slightly over-sized high resolution focal area you can detect a change in gaze direction before the user's focus exits the high resolution area.

Anyway that was ages ago and we did it with like three people, some duct tape and a GPU, so I expect that it should work really well on modern equipment if they've put the effort into it.


It is amazing how many inventions duck tape found its way into.


Foveated rendering very clearly works well with a dedicated connection, wiht predictable latency. My question was more about the latency spikes inherent in a ISM general use band combined with foveated rendering, which would make the effects of the latency spikes even worse.


Although thus isn’t directly related to the idea in the article, I’m reminded that one of the most effective hacks I’ve found for working with ChatGPT has been to attach screen shots of files rather than the files themselves. I’ve noticed the model will almost always pay attention to an image and pull relevant data out of it, but it requires a lot of detailed prompting to get it to reliably pay attention to text and pdf attachments instead of just hallucinating their contents.


Hmm. Yesterday I stuck a >100 page PDF into a Claude Project and asked Claude to reference a table in the middle of it (I gave page numbers) to generate machine readable text. I watched with some bafflement as Claude convinced itself that the PDF wasn’t a PDF, but then it managed to recover all on its own and generated 100% correct output. (Well, 100% correct in terms of reading the PDF - it did get a bit confused a few times following my instructions.)


This is probably because your provider is generating embeddings over the document to save money, and then simply running a vector search across it instead of fitting it all in context.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: