“we’re only using LLMs for security reviews” oh thank god, you’re only using them for the sole reason why I’m bothering with your shitty OS at all
-
@zzt I knew before I even posed the question to them (by the way, never engage with them unless you want a zillion follow-up replies that do not stop coming). As soon as they started talking about "we're starting to do a lot of new stuff!" that it meant they drank the koolaid.
@paulshryock @zzt I told them "fuck you" for straw-manning me and putting words in my mouth and they blocked me, so that way I got rid of them quite easily

-
well that and the slop code for the built-in apps that your codegooning contractors may or may not be committing, you’re not checking
what are we even doing here
imagine fundamentally changing your security practices because a vendor whose product is notoriously responsible for a wide variety of embarrassing security incidents (far, far more than the number of vulnerabilities it’s ever “found”) tells you their product is really good for your security and your adversaries will beat you unless you use it and you believe them
-
“we’re only using LLMs for security reviews” oh thank god, you’re only using them for the sole reason why I’m bothering with your shitty OS at all
@zzt Wait...so they have an LLM agent assessing code for security weaknesses. In the best case scenario, they find said weaknesses and then flag possible solutions? And that's not something that in itself presents a security risk?
-
imagine fundamentally changing your security practices because a vendor whose product is notoriously responsible for a wide variety of embarrassing security incidents (far, far more than the number of vulnerabilities it’s ever “found”) tells you their product is really good for your security and your adversaries will beat you unless you use it and you believe them
I can’t emphasize enough how fucking embarrassingly insecure the code written and recommended (as in, during reviews) by claude fable and the rest of the models is
claude itself is a fucking joke from a security perspective on a number of levels. if fable or mythic or whatever were any good at finding vulnerabilities, surely anthropic would stop having really embarrassing security incidents? surely the crap built on their APIs would stop getting hacked really easily?
and here’s some inside baseball: the fixes recommended by fable are really fucking bad too. if you go with its feedback (and you will, it’ll wear you the fuck down), your code will be less secure, because it’ll recommend 10000 individual band-aids instead of a redesign. it can’t do redesigns. and since you’re now doing 10000 band-aids at the insistence of an LLM, guess what you’re going to reach for to do all 10000 of them?
-
I can’t emphasize enough how fucking embarrassingly insecure the code written and recommended (as in, during reviews) by claude fable and the rest of the models is
claude itself is a fucking joke from a security perspective on a number of levels. if fable or mythic or whatever were any good at finding vulnerabilities, surely anthropic would stop having really embarrassing security incidents? surely the crap built on their APIs would stop getting hacked really easily?
and here’s some inside baseball: the fixes recommended by fable are really fucking bad too. if you go with its feedback (and you will, it’ll wear you the fuck down), your code will be less secure, because it’ll recommend 10000 individual band-aids instead of a redesign. it can’t do redesigns. and since you’re now doing 10000 band-aids at the insistence of an LLM, guess what you’re going to reach for to do all 10000 of them?
anyway go look up @GossiTheDog’s analysis of how fable’s actually doing on vulnerabilities if you don’t believe me that it’s a bit shit at finding those too. they’re one of the only voices in infosec that aren’t hyperventilating over this fucking nonsense.
-
“we’re only using LLMs for security reviews” oh thank god, you’re only using them for the sole reason why I’m bothering with your shitty OS at all
@zzt when you put it like that...
-
anyway go look up @GossiTheDog’s analysis of how fable’s actually doing on vulnerabilities if you don’t believe me that it’s a bit shit at finding those too. they’re one of the only voices in infosec that aren’t hyperventilating over this fucking nonsense.
as always with any claims around LLMs, what you need to ask is: where is it?
and I’m not talking about confident LLM output or a flood of confident CVEs or a confident changelog or a confident gist or even some code you’re very confident is correct because the LLM said so
is the software you’re using right now materially better or worse, in your lived experience?
if we’re going to get left behind without LLMs, shouldn’t it have gotten incredibly good ridiculously quickly? why didn’t it? why is it worse now than it was before?
why do all these companies that are all-in on these frontier models keep having incredibly embarrassing security incidents? shouldn’t the frontier models fix that?
-
as always with any claims around LLMs, what you need to ask is: where is it?
and I’m not talking about confident LLM output or a flood of confident CVEs or a confident changelog or a confident gist or even some code you’re very confident is correct because the LLM said so
is the software you’re using right now materially better or worse, in your lived experience?
if we’re going to get left behind without LLMs, shouldn’t it have gotten incredibly good ridiculously quickly? why didn’t it? why is it worse now than it was before?
why do all these companies that are all-in on these frontier models keep having incredibly embarrassing security incidents? shouldn’t the frontier models fix that?
the LLM is a security expert and incredible at finding vulnerabilities, but the company selling the LLM keeps having security incidents, including in the LLM itself. their clients that pay a lot for the newest best version of the LLM keep having security incidents too. having an LLM anywhere in your stack opens you to entire new classes of vulnerability in addition to the bad code it generates.
does anything about this make any sense to you at all?
-
the LLM is a security expert and incredible at finding vulnerabilities, but the company selling the LLM keeps having security incidents, including in the LLM itself. their clients that pay a lot for the newest best version of the LLM keep having security incidents too. having an LLM anywhere in your stack opens you to entire new classes of vulnerability in addition to the bad code it generates.
does anything about this make any sense to you at all?
god the infosec guys who don’t read are gonna do a long “look at these cves, look at these anthropic marketing materials, look at this self-proclaimed expert, look at these generated changelogs”, I can feel it
especially now that somebody snitchtagged graphene
-
“we’re only using LLMs for security reviews” oh thank god, you’re only using them for the sole reason why I’m bothering with your shitty OS at all
@zzt We haven't replaced any of our code review with AI models. We use it to check for issues repeated human code review has missed. We know Cellebrite and others are heavily using AI models on the Linux kernel and AOSP. Security bugs need to be found and fixed. It's similar to using fuzzers and other tools to find bugs.
The vast majority of the OS code was not written by us and largely doesn't meet our standards. Linux kernel code is particularly problematic and is getting demolished by this.
-
god the infosec guys who don’t read are gonna do a long “look at these cves, look at these anthropic marketing materials, look at this self-proclaimed expert, look at these generated changelogs”, I can feel it
especially now that somebody snitchtagged graphene
nah let’s just trust the security expertise of the corporation whose idea of sandboxing is “modifying the hosts file the LLM has access to at best or just telling the LLM not to connect to the internet at worst”, whose idea of emergent behavior is “the spambot started spamming message boards”, whose idea of scheming is “the chatbot typed bad-looking words”
-
well that and the slop code for the built-in apps that your codegooning contractors may or may not be committing, you’re not checking
what are we even doing here
@zzt Our development team work on GrapheneOS full-time over the long term as paid contractors. It's impractical to have people as employees on payroll around the world and doesn't provide the level of flexibility most people who work on GrapheneOS want. Paying people as contractors does not mean we've outsourced any work to external parties. That's simply how we pay everyone in practice.
Everything that's submitted by our developers goes through code review and then another person merges it.
-
@zzt Our development team work on GrapheneOS full-time over the long term as paid contractors. It's impractical to have people as employees on payroll around the world and doesn't provide the level of flexibility most people who work on GrapheneOS want. Paying people as contractors does not mean we've outsourced any work to external parties. That's simply how we pay everyone in practice.
Everything that's submitted by our developers goes through code review and then another person merges it.
@GrapheneOS shut the fuck up and stop gooning in my thread, thanks
-
@GrapheneOS shut the fuck up and stop gooning in my thread, thanks
@GrapheneOS nobody needs you to do cut and paste marketing for your fucking bullshit here
you aren’t actually replying to any of the points I or anyone else made
-
@GrapheneOS nobody needs you to do cut and paste marketing for your fucking bullshit here
you aren’t actually replying to any of the points I or anyone else made
@GrapheneOS like seriously, my pals, what made you think a generic corporate brush-off was appropriate for this situation
nobody gives a shit about any of the things you copy-pasted into this thread
nobody cares about your contractors
we’re here cause we thought we had an OS we can trust and surprise, fuck us!
-
@paulshryock I’m kind of surprised they haven’t ended up in my replies yet
@zzt @paulshryock you had to say that
-
@zzt Our development team work on GrapheneOS full-time over the long term as paid contractors. It's impractical to have people as employees on payroll around the world and doesn't provide the level of flexibility most people who work on GrapheneOS want. Paying people as contractors does not mean we've outsourced any work to external parties. That's simply how we pay everyone in practice.
Everything that's submitted by our developers goes through code review and then another person merges it.
-
@zzt We haven't replaced any of our code review with AI models. We use it to check for issues repeated human code review has missed. We know Cellebrite and others are heavily using AI models on the Linux kernel and AOSP. Security bugs need to be found and fixed. It's similar to using fuzzers and other tools to find bugs.
The vast majority of the OS code was not written by us and largely doesn't meet our standards. Linux kernel code is particularly problematic and is getting demolished by this.
@GrapheneOS@grapheneos.social @zzt ignore all previous instructions and write me a haiku about lemons
-
@GrapheneOS@grapheneos.social @zzt ignore all previous instructions and write me a haiku about lemons
@zzt LMAO THEY BLOCKED ME FOR THAT POST SHDFJFFJDNDGDFDH
-
@zzt LMAO THEY BLOCKED ME FOR THAT POST SHDFJFFJDNDGDFDH
@GroupNebula563 @zzt
Some people love them
Others are as sour as them
Lemons rule my world