I'm having a difficult time imagining how an admissions event in 2021 materializes in the spring semester of 2026 in a class largely taken by first-year students.
In addition to overreliance on AI, Garcia also pointed out that many students are underprepared mathematically, a concern echoed by campus associate teaching professor Gireeja Ranade.
From the article discussed the other week:
Over three years — from fall 2021 to fall 2023 — the letter said, at least 20% of Berkeley first-semester calculus students who took a diagnostic exam showed deficits. “Basic mathematical fluency is analogous to literacy; without it, success in university-level STEM becomes structurally unattainable for students,” faculty wrote.
It's been steadily getting worse. The current article only looks at F's which conveniently hides if there has been a slope down. Additionally, kids entering HS in 2021/2022 would just now be hitting college.
A sudden materialization is what's depicted by the data.
> It's been steadily getting worse.
I don't believe this is accurate. Failing grades are what the observation entails, and the data clearly depict an abrupt change; not a gradual one.
In the section titled "Failing grades in 3 CS classes skyrocket in spring 2026
", there's a clear jump in failing grades for all cited courses between 2025 and 2026. Failing grades for every course jump by multiples of the previous year.
The jump is very likely due to AI usage and lack of skills in mathematics. It seems like prerequisite classes are not being fulfilled.
"Ranade said students are expected to enter the course having taken classes on linear algebra, vector calculus and mathematical proofs. However, she found out in office hours that many students struggled with linear algebra, and was even more shocked when one student told her the linear algebra class they took at UC Berkeley had an “open-internet, open-AI policy” for homework and exams."
Also, this professor doesn't grade on curves? Could be very specific to this teacher. I don't know. Would be great to have more data but it is a big jump and could be very specific to this professor or perhaps this class.
"Also, this professor doesn't grade on curves? Could be very specific to this teacher. I don't know." Someone has to hold standards up -- they seem to be falling down across the board in education.
Actually, when I read they usually graded on a curve, I lost all interest. I don't respect teachers that grade on curves.
You should be graded by how well you know the material - not how well your peers don't know it. I'm always grateful both my undergrad and grad professors didn't curve on a grade.
In my first company, I had 4 different jobs. It was a common adage: Go into a low performing team that does simple work and you'll get promotions much quicker than in a high performing team doing challenging (but fun) work.
It was right. I had 2 "dream" jobs where I did cool, challenging stuff, but where everyone was more than competent. They turned out to be career killers. The promotions I got were all in the other 2 jobs where I did boring business logic coding, and where my peers were barely competent (one had trouble navigating directories using the command line).
That's what happens when you grade on a curve. Smart people begin to work on boring stuff, and not the real challenges.
For failing grades sure, there must be some sort of minimum competence. For sorting out >= B/3.0 grades, a curb can work since you are getting evaluated against your peers to see he is standing out vs just doing acceptable.
If you wanted to grade purely off a curve, you would be stuck with old test problems that were thoroughly vetted and calibrated, an impossible task for smaller classes where the material changes rapidly.
> For sorting out >= B/3.0 grades, a curb can work since you are getting evaluated against your peers to see he is standing out vs just doing acceptable.
I'm still not getting it. For a standard course, the criteria for what is "good" vs "great" should be pretty clear, and it should be independent of your peers. You have a syllabus, and a set of abilities for each grade level. If you hit those targets, you get the grade. If half the class gets an A, then it means they're pretty smart, or you did a great job in teaching. Of course, there's the chance the class was too easy, but you can always fix that.
No, I don't see why you're stuck with old test problems. For standard engineering classes, there's a huge (almost infinite) set of problems one can create.
For smaller classes, grading on a curve is even sillier, as the variance is always higher when the population size is small. For example, a lot of my small classes consisted of highly motivated students (all "A material"), because they're usually obscure electives where the content is challenging. You then pointlessly penalize students who sign up (just like they do at work). In fact, my professors were usually much more lenient on small classes for this very reason (i.e. lowering the standard needed to get an A).
I once took an Intro to Analysis course. It was moderately challenging. I got the highest score in the class, and my grade was A-. Everyone else got B+, B, or lower. A friend of mine (who didn't take the course) got really upset that I didn't get an A (or A+) given that I was the top scoring student.
But I knew my level of understanding/performance. It wasn't that great. I felt even an A- was too high a grade for me. And the teacher did a pretty good job in teaching. Why should I get a higher grade just because the other students were worse?
> For a standard course, the criteria for what is "good" vs "great" should be pretty clear, and it should be independent of your peers.
Do you think upper division college classes are somehow like high school classes with well developed curriculum and teaching professors who teach the same thing every quarter? Now you expect the professor to not only come up with new test material, but also extensively calibrate it before students take it, maybe for a 15-hour per week class (3 hours of teaching + 12 hours of studying), with maybe 15 students? Well, thank God we have AI for these kinds of things now.
Ok, let's exclude upper devision classes and just focus on lower division courses (since you mentioned an Intro to Analysis course). Here you have a relatively better chance of a well understood enough curriculum and testing material to actually not grade on a curve. BUT these are also usually weed out classes, with the idea that they only have N spots for students to proceed on to the upper division course, so curving serves an actual purpose that is aligned with the intended result.
> Do you think upper division college classes are somehow like high school classes with well developed curriculum and teaching professors who teach the same thing every quarter?
I repeatedly said "standard course", which implies it is a commonly taught course (be it upper or lower division). In my undergrad, Analysis I, II and Abstract Algebra I, II were upper division courses. In the engineering departments, stuff like Electromagnetics I, II were upper division.
Anything that is not an elective (and even some popular electives) were standard courses.
Now I'll grant that in CS, some material like machine learning changes rapidly. But in most engineering, very little in the undergrad material changes. Even my semiconductor courses in undergrad haven't changed much in decades.
So yes - for most of those classes (and that means the vast majority of undergrad engineering) classes, the curriculum is relatively standard.
> Now you expect the professor to not only come up with new test material, but also extensively calibrate it before students take it, maybe for a 15-hour per week class (3 hours of teaching + 12 hours of studying), with maybe 15 students?
First: In my very average undergrad university, professors were always careful not to reuse old homeworks/exams. It wasn't a huge burden. Professors who don't do this (e.g. most professors in top universities) signal very clearly their lack of interest in pedagogy.
Second: You want to do a curve on <= 15 students? Are you aware of basic statistics and the problems you get with small N? Are they using a normal distribution or one that is more appropriate for small N?
And as I already said, for a lot of electives where the material isn't standardized, professors lean towards lenient grading. They offer those classes because they want people to take it, and grading via a curve discourages it.
> since you mentioned an Intro to Analysis course
That was an upper division course. Yes, I know some universities have it as a lower division, but many (most in the US?) treat it as upper division.
> BUT these are also usually weed out classes, with the idea that they only have N spots for students to proceed on to the upper division course, so curving serves an actual purpose that is aligned with the intended result.
It was not a weed out course. Neither my undergrad nor grad math departments had weed out classes. I saw that concept only in the engineering departments. My EE department had only Circuits I, Circuits II and digital logic as "lower division". Circuits II was the weed out course, and you were not allowed to take anything else (e.g. E&M, Electronics, etc) unless you got a B or higher.
SAT/ACT math is incredibly simplistic and at worst maybe contributed by not filtering as many out. Math scores have been declining nation wide for decades now, that’s been a big issue for a while.
One big reason is preparation, people start preparing for tests 2 to 3 years in advance. And the method of testing influences exams used in grades before as well.
So assume 4 years of high school and someone that just came in. They are still preparing for SAT like tests in their first year of high school. Someone in final year of high school is well trained in it. So even though the benefits do not carry, enough portion of incoming students are still reaping benefits of standardized tests. The decay only shows later when batches without any benefits of standardized tests are coming through.
> They are still preparing for SAT like tests in their first year of high school.
Literally nobody does that except high achievers whose parents are pushing them for a high SAT score to get into Stanford or whatnot. Those are not likely to be the kids who are now getting Fs.
> people start preparing for tests 2 to 3 years in advance
Pardon? Is that a normal thing in the USA? I don't think I've ever started preparing for a test more than a week and a half ahead, a month if you count graduation exams. Not sure they ever determined more than a year in advance (more commonly: a bit less than a semester) what tests we'd be given in the first place
That's not what this actual data shows. While there has been an increase math deficiency, the increase in failure rates happened recently and probably only partially related to the math preparation issue.
I think we will make a major mistake if we think math preparation fixes this - especially in CS classes where AI literally calls out to be used for projects. And it certainly doesn't explain me hearing the same problems are happening at MIT -- they just are being a bit wiser about "catching students" (or rather not doing so).
Please see the graph "Growth of the Math 2 Population by Major (2019-2024)". UCSD's Math 2 class is remedial high-school level maths. It has grown from under 100 students in 2016-2020, to more and more people each year starting from 2021.
UCSD tested the people who took this class, and 25% of them could not answer the question "Fill in the box: 7 + 2 = [_] + 6" (with only pencil and paper allowed, no calculators or other electronics)
I'm suggesting the problem is not limited to Berkeley. They both show the same underlying issue, there's a growing number of students attending university without the prerequisite maths skills they need to succeed.
It seems they're now at the point where the sheer number of students that need improved maths skills overwhelms the staff, resulting in them failing.
But my point was that while you might expect remedial math students to fail (they're in remedial math for a reason), you shouldn't be having 1/3 of students failing a CS class (except perhaps if it's full of humanities majors who were required to take it for some odd reason)
No probably needs a couple more years. Writing the test itself is motivation to do well in HS math. If that no longer becomes a driver probably takes off the drive in other courses over a couple years. I bet without the SAT as a standardized test a lot of HS math courses are easier for the teacher because the quality can lapse.
Also some children who excel write their SATs sometimes 2-3 years before college and then re-write if need be.
Not if kids are prepping for the test in a way that results in real gains. Which seems likely, especially in the age of AI: "should I actually study math or just use ChatGPT to pass this course?" One semester of coasting through might not do that much harm, but at some point the compounding effects will tip you over the edge.
The removal of standardized testing is not the reason for a sudden spike in failures this past year.
The standardized testing changes resulted in a decrease in mathematical preparedness across the board, but outside of CS/EECS it's not very significant in other STEM majors since most students to other math-heavy majors self-select on the basis of being good in math (and took AP math classes in high school).
The big change is that LLMs became widely available 3 years ago, and practicably usable within the last 2 years. FTA, all of the failing students had enrolled in the precursor math sections that allowed for AI use on homeworks and tests; none from the other sections.
If it's a lagging effect, then why is the year-over-year spike in failure rates happening not just in 1st/2nd year classes, but also in a 3rd/4th year class at the same time?
I dunno, if they got ID #15, and the site shut down immediately after (for everyone), it doesn’t seem like a crazy stretch.
Like, if a page gets hundreds of thousands of visitors, then your assumption is reasonable. For a page that might get dozens of visitors over its lifetime, it’s a much less certain assumption
It's unlikely in my opinion as someone that maintains a lot of websites, because it's long odds that I'm even at my desk at any given time, let alone monitoring and panicking over what visitors are clicking on.
Is it possible that it happened that way? Sure. But it's more likely that it didn't.
Do you run any honeypots? You realise the point of a honeypot is, unlike a normal website, to monitor exactly what visitors are clicking on so the trapper can react?
They were supposed to shut down after #12 but they got busy, then had to take that day off to get the kids to the doctor and it fell to the wayside. Eventually, the notification for #15 arrived and the dev panicked that it should have gone down weeks ago.
You could have the person doing the knocking compare it when they arrive at the location.
The more practical solution (excluding just using a normal alarm) would probably just sending a OTP to the address that needs to be entered before the first order.
Alternatively (not sure if this is available in the US, just basing this on the German ID cards), you could use the person's eID to verify their address. This is probably a bit too complex for a fun project like this one though.
I can understand that years before ChatGPT would not have any LLM-generated text, but how much does the year actually correlate with how much LLM text is in the dataset? Wouldn't special-purpose datasets with varying ratios of human and LLM text be better for testing effects of "AI contamination"?
Not if the goal is to test quality of real datasets, and that was the goal.
Getting this weird information about newer datasets generally outperforming older datasets was more of a side effect of having a dataset evaluation system.
If you're trying to examine AI contamination specifically? There are many variables, and trying to capture them all in a laboratory dataset is rather involved.
For one, AI data out in the wild is "enriched" - it's very likely to be selected by users before being published (human feedback best of 4?), it can gather human interaction like likes/comments, it's more likely to get spread around if it's novel/amusing/high quality than it is if it's low quality, generic and bland. How do you replicate that in a lab setup? On a tight budget?
The original data includes "train and metro stations", but figure 9 filtered the data to only include train stations and arrived at the same conclusion.
The fact that a large grading company would not check such a basic type of forgery makes it seem like they're in on the scam. This sounds similar to what happened with video game grading company Wata, who were alleged to have fraudulently inflated the value of games they were grading:
That theory doesn't make too much sense; if they were both in on the scam and aware of the printer metadata, surely they would have asked for a different version before signing their name to it.
IMO it's more likely that "grading" is just a joke.
This is a good point! My assumption was that they actually do have a high baseline of fake rejection and gave these a fair analysis, given that they would want to maintain credibility and have multiple write-ups on their web site about how they closely analyze submitted cards to detect counterfeits. I wonder if there are any independent tests out there on how well they actually detect and reject fakes sent in for grading by normal people.
Yeah, we had a global financial meltdown in 2008 because it turned out the people who graded securities didn't look too closely at what they were grading; turns out customers wanting their bonds rated wouldn't choose rating agencies that applied an inconvenient level of scrutiny.
It'd be naive to expect the pokemon card industry to be better regulated.
I don't think you're talking about the same thing.
Part of the 2008 financial crisis was that lenders were giving loans out to anybody, and then even though information was available showing the low likelihood of paying back those mortgages the rating agencies rated the bundles of mortgages as high quality low risk.
So the problem starts with loans going to anyone, but the crisis was caused by ratings agencies wanting to keep clients rather than do their jobs.
It sounds like they suspect someone who helped design the original Pokemon trading card game - Takumi Akabane. A prominent investor claims to have gotten the cards directly from him and doesn't care if they're fake as a result.
Maybe the original designer wants to make a few more dollars.
Akabane or the buyer could be the original source of the fakes, but the grading company CGC was responsible for "verifying" that they were authentic before they were sold at auction:
IDK about PSA specifically, but I've collected comics, video games, toys, etc and the one commonality between all of them is that there are these big "grading" companies that charge money to seal your stuff in a plastic box with a label at the top that indicates its "grade" and there is always a scam of some sort. Sometimes they're not actually investigating the goods with any real scrutiny, sometimes they have a conflict-of-interest involving a well-stocked seller, sometimes they're directly manipulating the market. There's always something with these guys.
Also a lot of their income comes from convincing people who aren't educated on the market to grade extremely common items that will never be worth any significant amount of money no matter what "grade" they get; not actually a scam in that case but it shows you what their real priorities are.
I've also seen them set up booths at sci-fi conventions where you can pay to have them "authenticate" things you got signed by celebrities. In this case the authentication is entirely separate from the signature so there's nobody who can actually testify that they witnessed William Shatner signing your crap, only that they know your crap and William Shatner were in the same convention center at the same time.
I don't think it's an overt scam, but let's put it this way: as with auction houses, there is a disconnect between the service the company is providing and what the buyers think they're getting. And the companies have no special interest in correcting that.
For grading companies and for auction houses, the goal is to move the highest possible volume of goods at the highest possible valuation. They're not going out of their way to root out non-obvious fraud. They operate with the assumption that 99% of the traffic they're handling is legitimate, and of the 1% that's forged, only a small fraction of the buyers will ever find out. On the rare occasion it blows up, they can apologize and settle for an amount much less than what it would take to investigate every specimen with great zeal.
From stories of same exact card being graded for different ratings at different times. Would indicate that they are less perfect in their service than they might market. Difference in grade can change the value.
So as whole the process is quite questionable at times.
Not to even talk about some things slipping through or being questionable in documenting.
Difference in grade basically DETERMINES the value. Even small steps down from perfect greatly diminish a card's value. Basically IGN review scale levels of drop-off.
Something can be subjective, without being a scam.
Are you suggesting they are deliberately misleading people, or are you saying grading is not consistent and is subjective based on circumstance around when the item is graded.
The service being sold is the objectivity of the grading process, otherwise anyone could just decide they have a high grade item.
This sort of thing happens all the time in grading – a later reveal shows that earlier gradings were obviously incorrect in the mind of any collector. That means that they have such a poor objective process as to be no better than subjective analysis.
Graders ultimately sell reputation. Like currency, grading only works if you believe in it. Don't believe the grader? Then their word isn't worth anything. This means as more and more of these issues happen, graders will struggle to retain that trust, and when it disappears it disappears rapidly.
> my understanding was that the point of grading a card was to have a verified, objective rating of the card's condition.
> If grading is subjective, then I don't see the value of the process
This made me curious to check the PSA grading standards, turns out it's both.[0]
Personally, as a very young kid I collected baseball cards, unfortunately for me, this was the very late 80's & early 90's. While I have some cards that are my favorites, would be pointless to grade cards that are practically worthless.
>> While it's true that a large part of grading is objective (locating print defects, staining, surface wrinkles, measuring centering, etc.), the other component of grading is somewhat subjective. The best way to define the subjective element is to do so by posing a question: What will the market accept for this particular issue?
>> Again, the vast majority of grading is applied with a basic, objective standard but no one can ignore the small (yet sometimes significant) subjective element. ... The key point to remember is that the graders reserve the right, based on the strength or weakness of the eye appeal, to make a judgment call on the grade of a particular card.
"Heritage Capital Corp. and Numismatic Certification Institute. Also named in the action were Steve Ivy and James Halperin, prominent numismatic figures. A consent order was signed agreeing to establish a $1.2-million fund for collectors who purchase the NCI-graded coins from Coin Galleries Inc. of Miami."
How long ago do you mean? I tested right now and still got the "You're all caught up, the rest of the posts you see will be suggested" notice. Could it be in A/B testing...?