Lemmy account of natanox@chaos.social

  • 0 Posts
  • 110 Comments
Joined 2 years ago
cake
Cake day: October 7th, 2024

help-circle





  • There are parental controls, but it’s rare. FOSS projects usually aren’t very keen on building systems of authority, even parental ones.

    For smartphones check out /e/OS (Murena Smartphones). They added parental controls recently, not sure how good those are though.

    For computers your best bet’s probably a Linux with Gnome. They also added something recently.

    You might be able to combine both with a custom DNS service, surely there’s something out there for that usecase (NextDNS?). But that’s already getting pretty complex.



  • I mean, they “plan” to publish that info, right?

    I’d bet they used at least CommonCrawl (with it being the majority of data), arXiv, Wikipedia and the set containing all of Github. Probably not a lot of distillation.

    CommonCrawl is one of the reasons small websites and social instances get DDoS’ed by rules-ignoring AI crawlers. Wikipedia data basically always gets used without paying them. Github… well, it’s a prime example of how they broke millions of licenses.

    There’s also other stuff commonly used. Don’t get me started on the training sets for image generation. It’s beyond disgusting (and I’m not even talking just about theft and cultural destruction at this point).

    This stuff is a bottomless pit, and if those (F)OSS models really aim to be up at the top they’ll have to break every possible moral, ethical and legal rule just like everyone else. Even more so if they omit closed-source training sets.



  • Not just that, companies are using “smart” devices as AI scraper botnets to utilize private, basically unbanable IPs. There are only very few companies who I might believe them not doing it (AllenAI and maybe Mistral - tell me if I’m proven wrong pls). But OpenAI, Anthropic, fucking Google and Meta, they all treat your network as their personal internet extension. It’s reasonable to assume any “Smart” device with wifi access that isn’t FOSS most likely being your enemy.

    Our family Nextcloud already got taken down by OpenAI swarming it… overloaded and crashed php-fpm within a minute. At least one client blasting all endpoints still advertised themselves as OpenAI crawler.


  • Why does this remind me of the common argument made against democracy and for e.g. China… “it moves fast, less resistance, faster results”…

    And yet it’s still worse.

    The whole argument doesn’t hold up either. Your company - or the people behind it, if it’s new - just can’t be known as the usual profit-driven asshole. I mean, look at how the relations between the Community and Valve is. Nobody would argue Valve being perfect, but they’re known not to be the usual assholes and thereby cautiously welcome.

    If you or your company can’t deal with society being vigilant (especially in capitalism) then they’ve no place in said society. This vigilance is the only reason Linux still feels made for the user(advanced ones, but still). Fuck anyone who wants to destroy that with means of accelerationism and alike.

    DHH and Omarchy are just cultural cancer. What you correctly described is basically one aspect of how unfettered capitalism manipulates people and destroys healthy communities for the sake of growth.







  • Phew, to run such a model - even just 27B quantized - at a reasonable speed requires some serious GPU Power.

    Best thing I have is a 7800 XT 16gb (so ROCm it would be), don’t think that’s realistic and I’d rather invest the necessary amount of money for new GPUs into sth. more useful like a small electric vehicle than a slop cannon 🥴. Happy to hear actual useful things can be retained after the bubble finally crashes though. Oh please, let it crash tomorrow… together with the housing market…


  • Interesting. I tested ChatGPT and Mistral and oh my god, ChatGPT was abysmal. Hallucinated way harder. Mistral is okay if being used on highest settings.

    Did you create an account with Mistral? Without it they only expose a low-end model to the web, which isn’t properly declared last time I checked. Can be the free plan, I guess they just want to avoid their infrastructure being hammered.

    Personally all those Silicon Valley pricks can offer whatever, I don’t even have a Google account anymore and couldn’t be happier with that. Not saying Mistral is optimal, but at least they’re not part of the current US swamp. And I at least somewhat believe they actually read the damn GDPR. I’m totally fine with it not being the “cutting-edge”, especially given the cutting-edge cuts straight through culture and people’s lifes on purpose and with maximum brutality, and I seriously liked to see their image generator being way worse as well as Mistral refusing to open sites that banned it in the past (if they changed that behaviour please let me know). Not to mention their site isn’t even remotely as infected with trackers and other junk as the others, at least that’s what I’ve read (forgot to bookmark that hackers’ analysis I’ve read a year ago unfortunately).


  • Quick mention about Chatbots: In general they can be useful if you already have a grasp on something, so you can detect potential nonsense better. Additionally they can act like a drug especially on lonely people, and any company tries to sell them to you as awesome code generators. Yet Mistral AI already generated an answer to me telling me to dd my encrypted root hard drive for a speed test, even doubling down on it being safe.

    Not saying they can’t be useful for OP. Just that they’re like a handheld circular saw*, a powertool better only used by people who already know sufficient basics about wood.

    *where the salesman got rid of the safety handle and tells you to “just let it go on its own”.


  • The code quality apparently was abysmal and full of dirty hacks due to a mix of “beginner energy” and “go fast & break things” Silicom Valley mindset. The issues with bun (leading to the big Anthropic hack) were subtly blamed on the Zig language, even though Zig is completely fine; Bun’s code (which already was full of bad generated slop at that point, as well as code by programmers who were treated badly -> more bad code) was just crap.

    After Zig was blamed the Zig creator wrote a blogpost detailing all of it. Worth a read, certainly more human than the corposlop Anthropic announces.