Even if it works despite being AI slop that’s been developed in a mere 4 days (?) and somehow has all these contributors, I’m not sure I agree with the idea.
It’s cool that they went with Rust for it, but I would have much preferred if all that effort went towards improving GIMP and existing open formats.
If it works then I guess it’s the same end result at the end of the day, but it would suck if GIMP doesn’t end up being the one to put a nail on Photoshop’s coffin after all these years of serious development by humans.
Edit: I already see the kneejerk dislikes coming, I’m just interested in starting a conversation about the idea of it, I also dislike the fact that it’s AI slop as I already implied. I’d appreciate joining the conversation instead of dislike bombing with no argument.



The difference is the original source code, and similar source code that is incompatibly licensed, was in the LLM training data, making it not a clean room.
Thats not a difference. Other clean room projects, for example ReactOS, have source projects that are incompatibly licenced. Its the entire point of clean room reversing, to bypass copyright and licencing.
The training data is on the dirty room side, it doesnt get passed to the clean side.
The clean room concept is violated if raw source code is in the model generated, but to my knowledge its not. There is no way to pull out a chunk of the model, and say “This is a direct copy of main.c from xyz project”. At best, you can ask the LLM to guess what it should be, and it’ll regurgitate something statistically close. I just tried asking my local qwen3.8, and did produce a close summary of the nginx main file, but it wasn’t even close to a 1 to 1 copy.
The clean room would also be violated if you consider the training, model and inference as a singular entity, but that would be tough to argue when the inference can happen on one server, the resulting model isn’t functional in any way, its a big box of data, and the inference can happen on any GPU anywhere else in the world.
A human doing a clean room reimplementation is not supposed to have seen the original source code for something they are reimplementing. If that source code is in the training data for the LLM, that seems like a similar situation.
Except that they are separated. The clean side (model + implementors) doesnt have the original source code. The dirty side is the training room, which does have original source.
The model that you download from huggingface is basically the notebook passed between the two rooms, and as far as im aware, the model doesnt have the actual training data in it, just enough statistics to “hallucinate” something close.