You are not logged in.

#26 Yesterday 15:20:23

gromit
Administrator
From: Germany
Registered: 2024-02-10
Posts: 1,543
Website

Re: Policies on AI-generated code

No this meant an Arch Linux specific RFC (as in https://rfc.archlinux.page/), but it was not yet posted.

Offline

#27 Yesterday 19:40:37

Trilby
Inspector Parrot
Registered: 2011-11-29
Posts: 30,520
Website

Re: Policies on AI-generated code

KinuCyber wrote:

There's also the RFC on robots.txt ... which was discussed as one of the main component for AI control by content publishers

Unfortunately empirical tests have demonstrated that an exceedingly vast majority of the companies training AI models completely ignore robots.txt.  There have been other proposed methods working towards the same end goal, but most have proven to be a losing battle.  It would appear (a vast majority of) the companies training AI models have every intent of violating original author wishes to the extent that they can legally get away with it (which currently is to all extents).

Last edited by Trilby (Yesterday 19:41:02)


"UNIX is simple and coherent" - Dennis Ritchie; "GNU's Not Unix" - Richard Stallman

Offline

#28 Today 08:55:00

KinuCyber
Member
Registered: 2026-08-24
Posts: 4

Re: Policies on AI-generated code

gromit wrote:

No this meant an Arch Linux specific RFC (as in https://rfc.archlinux.page/), but it was not yet posted.

Ohhh, I see. I am pretty new here so still understanding stuff

You mentioned that there are talks about the rfc. If that's publicly available (discussion, forum, mailing list, etc) then may you kindly provide it?

Offline

#29 Today 11:08:58

KinuCyber
Member
Registered: 2026-08-24
Posts: 4

Re: Policies on AI-generated code

Trilby wrote:

Unfortunately empirical tests have demonstrated that an exceedingly vast majority of the companies training AI models completely ignore robots.txt

Not to be speaking in favour of such companies, but I find it hard for robots.txt to be a proper solution anyways. As stated in IETF's RFC 9969's section 2.1.2 ( https://www.rfc-editor.org/info/rfc9969 … on-2.1.2-3 & https://www.rfc-editor.org/info/rfc9969 … on-2.1.2-4 specifically ), the matter of respecting a person's preference for their content's usage to train AI is actually hard to implement. It should never be an excuse to steal another's work, but it's still something that requires looking deeper into the matter than just the binary approach of "use" or "not use" generative-AI.

So despite the popularity of robots.txt, it's not necessarily a viable approach all by itself unless supported by a wider (currently non-existent) framework.

Trilby wrote:

There have been other proposed methods working towards the same end goal, but most have proven to be a losing battle

That's one thing that, perhaps, we all can agree with.

Not that I am an ML Scientist, but my intuition says that the training frameworks need to implement a way to bake the credits of the dataset elements into the very neurons of the neural network (perhaps maybe dual-neurons side-by-side in same neural layer)

So that when the inference engine derives an answer, the answer itself can be referred back to the specific dataset element (content) it used and let us chain it all back to the content publishers involved. This should at least mitigate the issue with AI "stealing" the work of people who simply demand to be credited for their work.

Ofcourse, that's just an intuition by me. And even if viable, it's a long-shot and not something we can expect in just a year, especially when the environment for data collection has been hostile for decades so the currently available datasets themselves likely lack the necessary credits.


Nevertheless, generative-AI still uses our critical resources exuberantly so it's not really viable anyways in its current implementation. (Not to mention the current global economic model basically punishes anyone trying to use AI righteously.)

It's my personal opinion, and I'd like it to be challenged, that principle and pragmatism are not opposites. Instead, principled people need pragmatic mechanisms to actually enforce said principles. Otherwise, the principles remain symbolic. Your principle of enforcing policies on AI-generated code is venerable. Just that there's no one-size-fits-all conclusion and varying circumstances require appropriately varying approaches.


On a separate note,

Trilby wrote:

Unfortunately empirical tests have demonstrated that an exceedingly vast majority of the companies training AI models completely ignore robots.txt

You mentioned empirical tests. I believe you (especially since I am too witnessing all this myself) but may you still kindly provide links to the ones you've seen?

Offline

#30 Today 12:56:47

seth
Member
From: Won't reply 2 private help req
Registered: 2012-09-03
Posts: 77,825

Re: Policies on AI-generated code

So despite the popularity of robots.txt, it's not necessarily a viable approach all by itself unless supported by a wider (currently non-existent) framework.

It worked fine for 28 years as a proposed standard and two more years as formal standard - before the irresponsible dipshits bought enormous amounts of hardware resources with loaned money that the tax payers will have to pay back.

Offline

Board footer

Powered by FluxBB