Yet Another Rant About AI Security OSS Externality
IMPORTANT DISCLOSURE: This blog post is written in my personal capacity reflecting on the recent Apache Spark patch releases (3.5.9, 4.0.4, and 4.1.3). These are not the views of the Apache Software Foundation (ASF), the Apache Spark project, or any of my employers past or present. Both the ASF and the AI research lab have been provided an opportunity to comment on an earlier draft of this blog post and correct any information they see as incorrect and while many individuals have provided feedback in their personal capacity, there is not (as of yet) any official comment from either.
AI has fundamentally changed the volume of security reports in all projects, especially OSS projects where the barrier to analysis is even lower. We need to revisit how we handle security disclosures in the age of AI-driven pull requests and security auditing. Open Source Software (OSS) has long been for the community, and while some of us have jobs which give us time to work on open source, it's most often largely a volunteer effort for the public good, with the ASF focused on Software For the Public Good since it was incorporated in 1999. Security is of the utmost importance and requires greater care in collaboration to avoid accidental disclosure. For security teams, when working with OSS projects, please be flexible on release timelines when issues are not actively being exploited. If we have to move our release schedule to meet your planned disclosure date, it's more likely to cause us to ship more security issues rather than fewer.
While some labs have early access programs, in my experience they're not working. Despite there being an agreement for early access to one of the labs' "frontier security models", we only got access to a scan in August. While a team at that same lab, that we'll call Team X, is, seemingly, using their frontier models from at least February to report security bugs and issue deadlines for resolution and disclosure.

The old model of "here's a bug, you have ~90 days to fix your issues before we publish how to exploit your code" came about because software vendors, mostly closed source but some open too, were in the habit of ignoring security reports (or being openly hostile to them). Previously security reports took a human to make, and we'd get a few every month, but now with AI our security report inbox is looking more like the promotions tab in my Gmail. Additionally, more bug fixes are landing later in the release candidate cycle from AI driven PRs, so these previously reasonable hard security disclosure deadlines are making it more likely we ship worse security bugs while we try and play catch up. To that end, I want to call on all frontier AI labs to put more funding into mitigating the impact of their tools on the OSS projects which are vital to their success.

To put it in “corporate� the key asks from this entire post are fairly simple:
Frontier AI labs should provide real usable access to their frontier models for OSS
AI labs should be careful in teams using their models in house for security in advance of their own programs for “trusted� access.
More flexibility on disclosure dates is needed given the impacts of AI
A pony with attached coffee cup holder^ (and unlimited coffee) (jk)

The core of the issue remains that reviewer bandwidth is the scarcest thing in large open source projects. While tools like CodeRabbit are amazing (big shout-out to them for their OSS plans), they still don't capture everything. Thanks to the increasing frequency of attempted (and sometimes successful) malicious PRs and takeovers of OSS projects, we can't just leave it to the robots to argue amongst themselves. Maintainer burnout is a real security risk that we as an industry are not prepared to handle.

While these suggestions are only a starting point, the big AI labs can do a number of things to mitigate some of the downsides. One is providing usable non-time locked free accounts to the people whose software they used. For some of the bigger projects, the AI labs have staff who are able to do things like manage releases, but they're often too busy working on internal issues to have the time to work upstream. I think giving engineers something akin to Google's "20% time" (which some jokingly called 120% time) to give back to the projects they're using could help.
I want to double down on the "usable" part. While Anthropic has done better than any of the other AI labs at trying to provide free tooling for OSS developers, their effort still doesn’t do enough. In trying to use my Anthropic provided OSS maintainer account on a recent Spark security issue I tripped its safety guards and eventually ran out of tokens*. In practice, I've found paid inference on open weight models** to be more effective at addressing security issues. Most OSS maintainers, including myself, don't yet have usable direct free access to the frontier security models, meaning we're pushed into a reactive mode rather than pro-active.

For me these issues came into focus during the recent ASF Spark 3 and 4 patch releases, available everywhere fine bits can be downloaded. Most impactful was a security issue reported by Team X where because of a combination of a lack of communication in the initial disclosure and lack of flexibility on the surprise deadline once the release managers became aware, we had to scramble for extra reviewers on complex separate security issues.
Team X reported security issues in the Apache Spark project around February 25th of this year. This report was (thankfully) accompanied by a human-reviewed proposed patch and no specific deadline in the initial report, which to be fair is better than we would get with a frontier model alone. So far, so good.
However, out of band, Team X then gave a vendor who distributes and contributes to Apache Spark (not my current employer) 90 days to fix this, or Team X would disclose the issue publicly. While the vendor requested an initial extension to line up with the planned OSS release schedule, which Team X granted, the deadline was not communicated to the broader OSS group, including myself, until much later.
While working on the most recent releases, we got roughly 30 post-first-line filter security reports in the week before the first patch release candidate, a few of which turned out to be actionable. Part way through the release cycle, I became aware of the deadline from Team X and reached out to the head of security at the lab to further extend the deadline so we could address the outstanding security issues. Some of these issues were blocked on knowledge of specialized libraries I personally don't have. Initially Team X was unwilling to extend the deadline any further, but when I explained how our 72-hour release voting normally works, they were willing to grant one extra week if the vote failed. During this negotiation, the amazing ASF security folks found a reviewer for the complex issue I was wrangling. We were able to get known issues fixed and the release candidate vote thankfully passed, but that outcome was far from certain.
If we had not been able to find reviewers for the encryption fix, we would have had to ship a partial release with a known broken encryption issue. At the end of the day, Team X has still not published the disclosures as of the end of September, strongly suggesting that they could have given us more time. As a release manager, this is of course frustrating and had personal impacts on my vacation coffee consumption being more necessary rather than enjoyed, but the impacts could have been much worse (not just missing out on limoncello).
I am not particularly anti-AI as far as things go. Having worked on making AI tools to help people appeal health insurance denials, I think there are potential positive cases where AI can help people.
I'd also love to give a shout-out to my point of contact at the frontier AI lab for helping unravel the game of telephone and getting me that extra week.
I don't think this problem is unique to Team X, the process is broken and compounding the existing reviewer constraint. Reviewer bandwidth has almost always been the biggest bottleneck in OSS projects, and the decreasing cost to generate code will continue to push against this bottleneck. When issues are discovered in OSS projects that are not actively being exploited, requiring releases from the OSS project with a fixed timeline is counterproductive. Release cycles are likely going to keep getting longer (see how big the latest Linux RCs are), and we should allow time for this. For the humans driving the robots, when there is a public interest-reason like an active exploit, not just a conference or PR opportunity, for providing a disclosure deadline, make sure to include that in your initial report whenever possible.
I'd like to end by reiterating my request to all of the frontier AI labs: please offer truly usable free plans for OSS developers, be flexible with your disclosure deadlines when no active exploits exist. In addition, there are some amazing engineers at these labs, if you give your staff time to participate in OSS maintenance tasks it could make a big difference. There's a strong history of industry and OSS working together in both ASF and other projects from Spark to PyTorch to the Linux kernel. That and a pony with a coffee cup holder and a coffee, but I'll take more tokens in place of a pony (I still want the coffee though :p).
If you're coming to Glasgow for Community over Code, I'll be around and looking to catch up with other OSS folks for their views on how we can try and make the externalities from AI less terrible, including chatting with folks at the responsible AI round table, and talking about how to make PySpark even faster. For SF based folks I'll be at QConSF in November talking on using AI powered transpilation to make UDFs faster. And for folks regardless of location I'm going to be doing a virtual talk end of October on dealing with AI security reports (I'll post on LI and blog when the registration link is available).
P.S.
I am well aware this is a bit of a “do as I say not as I did� situation as I did report some Yahoo! Mail security vulnerabilities purely through side channels back in University almost two decades ago. But I was young and did not have the resources of a multi-billion dollar company behind me, and I would argue the security implications of unencrypted e-mail passwords are a bit different than data processing software with schema inference driven class instantiation vulnerabilities.
* I try to use the account for other OSS work, not just security. It is more successful for non-security work.
** I don't currently have enough vRam to run these on my own hardware, but if anyone has a few hundred spare RTX5090s or DGXs and a few 20a circuits give me a call :p
^ Did you know that's a real thing? I mean not the unlimited coffee, but they make coffee cup holders for saddles. In practice since I live in a city and do not have space for a pony let alone the skills, I ride a Tiger 900 (without coffee cup holder) and an LX150 (with coffee cup holder). If the AI labs are reading this: coffee cup holder for motorcycle with unlimited coffee is an acceptable substitute 😛
Frontier AI labs should provide real usable access to their frontier models for OSS
AI labs should be careful in teams using their models in house for security in advance of their own programs for “trusted� access.
More flexibility on disclosure dates is needed given the impacts of AI
A pony with attached coffee cup holder^ (and unlimited coffee) (jk)
I'd like to end by reiterating my request to all of the frontier AI labs: please offer truly usable free plans for OSS developers, be flexible with your disclosure deadlines when no active exploits exist. In addition, there are some amazing engineers at these labs, if you give your staff time to participate in OSS maintenance tasks it could make a big difference. There's a strong history of industry and OSS working together in both ASF and other projects from Spark to PyTorch to the Linux kernel. That and a pony with a coffee cup holder and a coffee, but I'll take more tokens in place of a pony (I still want the coffee though :p).





























