In the space of 1 week, a second open-source Chinese AI model equals the best investors are pouring tens of billions of dollars into.

schizoidman@lemm.ee · 2 months ago

In the space of 1 week, a second open-source Chinese AI model equals the best investors are pouring tens of billions of dollars into.

catloaf@lemm.ee · 2 months ago

Let’s give it a whirl!

welp

not_amm · 2 months ago

It works in Spanish, in English it throws an error before answering about Tiananmen as you show 🤫

catloaf@lemm.ee · 2 months ago

Interesting. I tried Chinese and it also throws an error. Looks like it was a manual thing in only some languages.

Karna · 2 months ago

Someone gagged the AI before it could complete that sentence 😜

Scrubbles@poptalk.scrubbles.tech · 2 months ago

I wonder if that’s a UI block like if it’s mentioned then throw error, or if the model itself has a block in there. From this, it looks like it’s baked in and of course they haven’t poisoned it

catloaf@lemm.ee · edit-2 2 months ago

I’m guessing it’s in the output handler, not the UI exactly. I don’t think you can edit models like that, and the fact that it knows about it at all means they didn’t whitewash the training data set. But my knowledge is limited. In their place, I would probably have included “don’t talk about tiananmen square” in the initialization rules. But failing that, I would have added something in the output processor to check for forbidden knowledge and throw an exception.

Still, it’s strange that it got the words out before dying.

Scrubbles@poptalk.scrubbles.tech · 2 months ago

Yeah agreed, I’m more surprised they didn’t scrub every reference to it on the training set like you said that it’s in the model at all is surprising. I may try to run it myself and see what it does with the same question

technocrit@lemmy.dbzer0.com · edit-2 2 months ago

These data processing apps are just apps hooked up to big computers with big data. It’s no surprise when rich people buy a bunch of computers and data then run an app on it. It’s much more surprising that people are hyped into believing this is somehow important. “It took one week to copy an app and load the data!!! Wow!!!”

Nobody cares when “China” writes a decent word processing app or whatever, nor should they.