The mathematical community was very competitive in its early years, but in the last 70 to 100 years, it has been generally less competitive and very collegial. The community was in a good place, and progress has been very good. In a few cases when competitiveness was ramped up, it lead to bad behaviour and destructive fights. Few would like to return to those competitive years.
I think both the competitiveness and and lack of problems will mean fewer people will do mathematics. I think collaboration will always be there, but the community as a whole will be smaller and weaker.
It's not so clear that AI is an amplifier.. The paper has some fascinating analysis on this topic:
"At the other end of the distribution, AI students who spend more than 65 minutes on their homework receive homework and exam scores similar to those of non-AI students, suggesting that these students do not use generative AI for homework assignments. However, this group consists entirely of students who adopted generative AI no more than Öve months. Six months after adoption, no AI student spends more than 65 minutes completing their homework (see Figure A5). This is consistent with the gradual process of learning how to use AI tools. It also suggests that AI crowds out the highest level of e§ort."
"Interestingly, in the range of 50-65 minutes, the median and the interquartile range
of exam scores of AI and non-AI students are similar. This implies that, in the range
where AI students and non-AI students have overlapping homework times, students
who spend the same amount of time completing homework on average receive similar
exam scores."
"This pattern shows that students who spend the same amount of time on homework
learn similarly, with or without generative AI. In other words, generative AI reduces
time spent learning for the majority of AI students but not learning efficiency for those who spend the same time studying as the non-AI students."
I wondered the same. ff and fi are common ligature letter pairs, maybe the original is in an encoding where those are represented with separate characters, which then correspond to § and Ö in whatever encoding we find here?
PDF is a render-only format with optional hidden metadata for semantics.
The PDF file they created knows how to paint the ligatures they used, like "fi" and "ff", but doesn't contain the `/ToUnicode` metadata for what those symbols meant.
The generated PDF happened to put the "fi" ligature at an index that is normally used "Ö", and didn't provide more information.
It was copied from a pdf, maybe LaTeX generated. All the ligatures are wrong -- not sure where the problem lies, the pdf itself, pdf reader, or the clipboard on Linux.
Depends who's using it. Like many tools the force multiplier depends on the operator.
I'm confident it's an amplifier for people who know how learning works and already do a lot of it, successfully. However the level of "learning fluency" I'm talking about isn't reached for many until late college or grad school, and sometimes not at all. So I'm not surprised by the quoted results for 12-18 year olds.
Fair question. In theory you could do an experiment where the subjects were grad students or professors, or top performing college students. Give them some fixed amount of time to understand some new subject on which they'll be tested, only 1 group has access to an LLM with appropriate context for the learning task, etc. That's just one idea...
This experimental setup would only tell you whether this demographic (grad students or professors, or top performing college students) benefits from access to LLMs for acquiring a new knowledge or skill.
It would not prove that "the results depend on the operator". To prove this (and figure out the traits that make a proficient operator) you would either need a massive dataset to which you'd apply some machine learning to figure out the correlations, or at least you'd need a hypothesis on the traits you want to test.
Do left-handed people perform better with LLMs? Analytical thinkers? Dyslexic people?
It's not enough to just claim "some individuals will perform better with these tools" without any notion of who does, and whether it's possible to become one of these people, and what is the expected gain in this group.
Otherwise, sure, we could sell unsecured chainsaws as a tool for making ice sculptures and observe that "some people" are indeed more proficient at not getting their face cut off ; but if the tool is being marketed to schools, such a vague statement is not helpful in making a case for it.
Yes, I'm using "high-performing college students" or "phd students" or whatever as a proxy for "people skilled at learning". It's not perfect but I think it'd be good enough. You could do lots of refinements on the idea. You could try to do some kind of "pre-test" that more directly tested "meta-learning skills".
My only point was I think it's possible to do. I personally think it's pretty clear. And there's nothing special about AI. In rural places, you could replace it with "access to books" or "access to a good tutor" and so on.
I also think you'd see the same effect with "motivation" (if you could test it), with performance on standardized tests, and with grades. AI will boost the better (by these metrics) students more.
My understanding of this study is that it disproves your assumption: "The negative learning effects are larger for students with higher initial achievement."
It seems that the students who had the most skill and motivation from the start, have the most to lose.
And people who want to learn. Most teenagers lack agency in their studies. They aren't in high school because they love it but because they have no choice.
> Sometimes things are just common sense pure and simple.
The discussion is about "AI", so common sense is out the window. These people's professional reputations depend on addict-level "AI" usage remaining socially acceptable.
well most university exams are designed to measure how much you study. so we didn't really need a study to tell us, "Exams continue to measure what they are designed to measure."
they're not designed to measure general aptitude, or function as admissions criteria, or screen for job applications, or any other numerous things they are used for.
there can be many questions of pedagogy. one of them is, what do our exams measure and how do we use them? professors who say, "My exam is designed to measure who studies, not be used for all these other purposes that they are actually used for" - I don't buy it. It's the same as late night comedians saying they are not responsible for solutions, even when spending 90% of their air time making political jokes.
THIS is the pedagogical issue, that pedagogy has NEVER caught up with the scope of responsibilities. This is acute in STEM - I mean, the humanities departments are generally pretty well run, all things considered, in this regard. Generative AI is accelerating that pre-existing crisis.
The fact that exam scores are correlated with how much you study is not the same as exams only reflect how much you study. Two students who study the same amount could have very different exam scores. The reason that there still is a strong correlation between exam score and time of study is because if all other things being equal, students who studies more have higher exam scores.
i'm not saying they reflect how much you study. they reflect a lot of things, including that. but you ask the people who write the tests, they're going to say, how much you know or how much you study, but nonetheless, they are limited. i agree with you. that's part of my point.
let's imagine a different study. we instead compare AI-users and non-users on a Wechsler (IQ-adjacent) test.
overall, it would be surprising if AI usage impacted your Wechsler scores. someone has done this study and the impact is quite quite small. BUT. do we care? We don't use Wechsler scores for admissions, we don't use them for jobs, we don't use them for... are you getting it now? A Wechsler family test is measuring something real, just like a university exam measures something. But what do we USE them for? Wechsler and a typical university exam are, in some senses, EQUALLY vague in terms of their fitness for purpose for answering a question like, "should we hire this guy?"
Like there is an association between IQ and earnings but it is actually surprisingly small! There is an association with math education and earnings and it is also surprisingly small. And consider how many people get by just fine without using a single piece of math education once they have finished school - like what if maximizing your earnings isn't all that it is about? Are you getting it now?
The issue isn't the AI usage. I can find tests that are immune to AI usage. The issue is using tests for things that they are not designed for. We pick and choose, for some subtle but nonetheless pervasive cultural reasons, which tests we use for which purpose, and very frequently, not because they are calibrated for the chosen purpose. This is coming from someone who scores very well on all these tests, and have kids, so I have a very strong incentive to buy into the status quo, and I'm telling you: academic testing has been fucked up for a long, long time.
Education is a complex topic. I don't claim to understand it, nor do I think we can sort this out in a hacker news thread. What I reject is the simplistic claims like "exams don't mean anything". There are certainly exams that are badly designed, but there are also exams that are well designed. A well designed exam can look very badly to different people, based on what they know and where they come from. It's a bad idea to think exam scores as the single metric of education quality, it is equally bad to reject them completely, because chances are any alternative measurement people come up with are going to be worse.
I think this deviates from the original topic. The article's intent is to argue that AI helps with learning itself, while you question whether this sample can serve as a reference. But what if the target of this sample is a group of serious college students? The article questions AI's enhancement of learning ability, not whether that exam has a real purpose.
You're right that academic exams also measure something along the lines of instruction following / obedience / willingness to jump over hoops for no good reason etc. and that's often a good signal for most kinds of jobs.
I just started teaching undergrad CS courses after ~20 years of various non-academic jobs and it took me about two semesters to realize that almost nothing matters except how I assess students.
Two weeks of ADHD-fueled research later, I concluded that academia is actively resistant to implementing assessment reform because it would expose the utter pointlessness of most of what happens in university classrooms.
The reality is that we have no idea what most university exams measure because they are ad hoc, written by amateurs (yes, most professors are untrained in pedagogical methods) with zero psychometric validity analysis.
> almost nothing matters except how I assess students
This is a reasonable opinion after two semesters.
IMO, with more experience you should have modified your stance.
If all classes were pointless, students would leave at the same point they entered. Since we observe that is not the case, something is happening in class. Do more of that.
The explicitly stated goal of most courses is mastery of a well-defined set of domain-specific knowledge and skills. The only way to objectively determine if the course achieved that goal for a given student is comprehensive assessment. Bad assessment design, much like bad experiment design in science, is worse than useless—it actively misleads and confuses.
A well-designed, repeatable, reliable assessment is a prerequisite before one can even consider the effects of particular pedagogical methods inside or outside of the classroom itself.
at least in my experience in university - i didn't really ask this question, since it is obvious to me, but some students have asked it during lecture, or some instructors have volunteered the answer ahead of time - if you ask how to perform better on the exam, usually the instructors say, "here's what you should study." they never say, "know more." the thing i am talking about is consistent with the paper. really, your takeaway should be, exams can't see how much you know!
Because "know more" isn't actionable. Knowing more is achieved by studying but not necessarily more time spent studying, but well spent effort. Staring at the page for hours and saying "I don't understand" doesn't help. Solve exercise problems, explain the material to fellow students, discuss it with them, make mind maps, bullet point summaries, work through derivations step by step, etc. There are many techniques.
At the end of the day though what matters is what you know. Furthermore, if it's a serious subject, it shouldn't matter whether you learned it from this teacher or from another school and teacher, as long as your knowledge is correct. Knowing the idiosyncracies of this particular teacher should not factor into the grade. A serious subject can be learned on one continent and examined on another. Bullshit courses are all about learning pet peeves and hobby horses of a particular teacher.
I would usually say something along the lines of “everything we covered is on the table” or “everything we covered since the last exam is on the table” depending on the nature of the test. That’s the same message as “know more” but I think it sounds politer.
I never studied in university and yet I acheived good exam scores. If you understand a topic and ave a reasonable memory and ability to apply my our understanding not much studying is required in my experience. If you don't understand the big picture then you got to laboriously keep track of and manipulate a bunch of disparate pieces.
Social stigma against just juices people to stick to socially desirable answers; no doubt a whole bunch who self reported as non-AI users actually used AI
Sounds like the opposite of the social stigma in tech companies, where most people are under pressure to self-report as AI users when they are actually just using normal methods to get the same result (and spending the time saved however they like) (nothing against genAI for making work take more time, it's still good for business as long as the extra work can be moved around)
It would be good to see the effect of access to AI during preparation on those who previously achieved top 10% points in exams of similar topics before. I suspect they would benefit further.
My apologies for coming off as over-enthusiastic, I am currently obsessed with this study. Here is another quote:
"The negative learning effects are larger for students with higher initial achievement. The differences in the estimated full (6-10 month average) effects are substantial, with a 50% gap between the most negative effect (-24 percent) for the highest tercile and the least negative (-16 percent) for the lowest tercile. "
Not top 10% as you asked, but the closest to what you asked. My working hypothesis is that top performance is highly correlated with willingness to work hard, and AI decreases the motivation to work hard.
Interesting. Top academic performance is mostly correlated with conscientiousness (willingness to work hard and keeping track of things) and intelligence. And I'd add motivation and interest to that too.
I think if you take a physics class where the student is intelligent and intrinsically motivated through their own interest (I admit this is rare) then AI probably helps.
I think if they are intelligent, intrinsically motivated, and willing to pursue knowledge beyond what the class requires, it probably helps.
That last part is key. Intrinsic motivation doesn't mean you pursue it outside normal bounds. My kid loves soccer, its her second favorite thing in the world, she has an absurdly high tolerance for physical discomfort while playing, but she doesn't play it at home. There's other things she's rather do, such as play with her toys.
When you move the bar to something even less interesting to most kids like science, you're going to have a pretty huge falloff. You're basically selecting for kids who choose to do it in their spare time. I know a lot of smart kids (I run a boyscout troop, my wife a girlscout troop, both with lots of high achievers), and none of them do this.
I think a substantial minority of students like subjects beyond what the class requires. They go to extracurriculars to learn more interesting math and science, or are history buffs who read about it on their own time etc. These are the people who as adults move things forward and solve novel challenges so it is important to equip them well.
Personally in my own work it's pretty obvious how it's an amplifier. You waste far less time trying to find an answer to something that confuses you.
You might say "it's good to learn research skills" and that's true to an extent, but tutors have always made people better students. And AI is a tutor you can message at any time, day or night, for free.
"AI is a tutor you can message at any time, day or night, for free"
which makes it not a tutor. if it is true that tutors have always made people better students, then those tutors are definitely imposing some limits on how many answers they give you and requiring you to do some thinking. I think you could have premised your same argument by, "copying a smart kid's answers has always made people better students...."
That's why you see the split in outcomes. The fraction of kids who value self improvement are going to get smarter, and the ones who just want to slide by will fall behind.
Nothing stopping kids from asking AI to give you responses like this, it is more than capable. There is always going to be the temptation to take short cuts though.
I wouldn’t be surprised if you did and still had the same issue, but if you relied on the training data recall alone, you definitely shortchanged yourself.
I’ve had decent results with tasking a model to run pre-research and summarization as a checklist, then have a second instance of the model(s) review and then develop my learning plan learn, versus the times I just asked Claude or ChatGPT to explain something to me “from memory.”
The knowledge AI has is vast. Communicating it to you is slow and narrow baud. The chat interface itself as a medium is lacking and will need to be replaced without another medium.
Finding an answer isn't always an objective, though. I agree it often has been too much in baseline schooling.
First, there's a certain amount of baseline knowledge we'd like students to possess. Without a certain prerequisite amount of underlying information committed to memory, it gets far more difficult to achieve fluency in a topic.
But more than that, there's other skills we're trying to build: frustration tolerance, processing contradictory information, disciplined problem solving. You only really develop these skills through productive struggle. If you find a way to shortcut the productive struggle, students truggle.
> but tutors have always made people better students.
Sure. Bloom showed us that students taught with a combination of tutorial and mastery methods, one-on-one, outperform students in a normal classroom by roughly 2 sigma.
The paradox has always been-- why hasn't technology unlocked these gains for students in normal classrooms? If we could boost everyone's performance by this amount, it would be huge for society-- but society can't afford to teach everyone with tutorial methods.
Since the 1970s, we've invested in edtech towards trying to make this happen, but most of it has actually had net-negative effects as best as we can measure. AI, so far, looks to be much worse.
I think part of the answer is that a big part of what makes a conventional classroom work are social pressures. So far, it looks like AI (and edtech in general) does more to dismantle conventional pedagogy and to break down the social fabric of the classroom, than it has improved differentiation or unlocked this tutorial effect more broadly.
the difference in the vast majority of cases is that people go to a tutor with an intentional stance of cultivating a particular kind of practice, not merely to get an answer to a question. It's not technically impossible to do this with a chatbot but far less likely given that, unlike an even half decent tutor, a chatbot will never cultivate that attitude in you. The elimination of that friction is exactly why people talk to bots.
The good thing about a tutor is that he or she isn't always available and knows they won't be in the future, so they instill good habits and independence in you.
> a chatbot will never cultivate that attitude in you
Hence the whole "AI is basically an amplifier of bad and good" argument made earlier. The ones who do this because they want to understand, don't need to cultivate this at all, it naturally happens with chatbots. They don't just ask for the answer to a question, but then dig into why it's like that and what not. But the ones that don't care, now have to do even less to get the fast answer without understanding.
Even if we grant that education is to prepare for employment, which a lot of people disagree with, it does not follow education should do the same thing as what's done in employment.
Tenured professors do often fail large swathes of the class, and it's not hard to stand their ground because academic freedom is still very important in universities. This is not generally true for non-tenured and adjunct professors, but for a different reason -- their job review rely on a large part on student feedback forms, and failing students are not happy students.
The idea that if only all professors stood their ground then somehow students will be motivated to study doesn't pan out in practice, though. There is already a significant number of students who are perpetually struggling. They are missing basic prerequisites, and instead of catching up on them, they repeated try and fail at learning the same materials, passing only when they got a lenient instructor. The problem compounds because failing brings helplessness and exacerbates their mental issues, which brings more failing. The university cannot sit on their high ground and watch these students struggle, especially if their number reaches a critical mass.
It's really tiring that LLM fans will claim every progress as breakthrough and go into fantasy mode on what they can do afterwards.
This is a really good example of how to use the current capabilities of LLM to help research. The gist is that they turned math problems into problems for coding agents. This uses the current capabilities of LLM very well and should find more uses in other fields. I suspect the Alpha evolve system probably also has improvements over existing agents as well. AI is making steady and impressive process every year. But it's not helpful for either the proponents or the skeptics to exaggerate their capabilities.
It's really tiring that LLM skeptics will always talk about LLM fans every time AI comes up to strawman AI and satisfy their fragile fantasy world where everything is the sign of an AI bubble.
But, yes this is a good way to use LLMs. Just like many other mundane and not news-worthy ways that LLMs are used today. The existence of fans doesn't require a denouncement of said fans at every turn.
I am criticizing how AI progress is reported and discussed -- given how important this development is, accurate communication is even more important for the discussion.
I think you inferring my motivation for the rant and creating a strawman yourself.
I do agree that directing my rant at the generic "fans" is not productive. The article Tao wrote was a good example of communicating the result. I should direct my criticism at specific instances of bad communication, but not the general "fans".
One could say the same about these kinds of comments. If you don't like the content, simply don't read it?
And to add something constructive: the timeframes for enjoying a hype cycle differ from person to person. If you are on top of things, it might be tiring, but there are still many people out there, who haven't made the connection between, in this case, LLMs and mathematics. Inspiring some people to work on this may be beneficial in the long run.
GP didn’t say they didn’t like it. They criticized it. These things are not the same.
Discussions critical of anything are important to true advancement of a field. Otherwise, we get a Theranos that hangs around longer and does even more damage.
I don't think you read the comment you replied to correctly. He praised the article and approach therein, contrasting it to the LLM hype cycle, where effusive praise is met with harsh scorn, both sides often completely forgetting the reality in the argument.
It's interesting that while Bourbaki had a large influence on modern mathematics, very few people read their books (at least among the people I know). In a sense, their project of producing a definitive exposition for a large part of mathematics has failed. I wonder whether it's because different branches of mathematics have their unique personalities, and therefore the attempt to provide a unified point of view are bound to fail.
I read once that the general attitude of the group was that their publications were not meant to be widely read, but just to provide the foundation for better expository work.
I also heard that part of the bad reputation that Bourbaki got was due to their being used in graduate education, despite warnings that they weren't suitable. In the 1950s/60s, there was a lack of good graduate texts. Of course, then Serge Lang came along...
Yes, Whitaker & Watson (analysis), Hardy and Wright (number theory), Dieudonne (analysis and he was literally a Bourbaki member), heck, Euclid's Elements; Gauss Disquisitiones, etc. Bourbaki is more of a monument. Writing it was necessary, but for readers it suffices to know that it is there ;).
while it's certainly not read by most mathematicians, Bourbaki (especially set theory & general topology) are still quite often read by mathematicians in training I believe.
I was applying a unfair standard to them of course. Every field has a few classics that last a long time, but most old books are not read. But I think Bourbaki maybe had grand ambitions that were eventually unrealized. My theory is that the presentation of mathematics is not based on unifying principles, but rather on the collective taste of mathematicians. So what end up being the most popular books is based on how the collective taste evolve.
they provided a unified point of view by explaining it all in terms of sets
ultimately they failed because they wrote such that it didn't matter if other people understood. it's a style that is only intelligible if you already know (from some other experience) what they are describing.
Bourbaki is known for their "definition-theorem-proof" style, which for a while influenced a lot of mathematical writing. It makes the logic of the presentation easy to follow. The proofs are complete and fairly clear. The logical order within books and in the series of books as a whole is also pretty good - if you read pages 1 through n in the books, you have the prerequisites to read a proof on page n + 1. There is a good index, a table of notation, exercises (at the back, not by section), and a table of contents (at the back, since the books are in French).
They probably originated the "dangerous bend" symbol (a Z-shaped curve in the margin) to indicate a tricky or subtle point.
They're pretty good as references (to look up the proof of a result, or read about single topic).
On the negative side:
There is little exposition in the sense of motivation for what is presented, or applications.
I'm looking at "Algèbre - Chapitre 10 - Algèbre homologique" (the only Bourbaki I own). In the introduction, they say:
"Le mode d'exposition suivi est axiomatique et procède le plus souvent du général au particulier."
"L'utilité de certaines considérations n'apparaitra donc au lecteur qu'à la lecture de chapitres ultérieurs, à moins qu'il ne possède déjà des connaissances assez èntendues."
Thus, you won't find applications, or many examples - just definition-theorem-proof.
It's assumed you know why you're reading the material, and so don't need to be told.
This particular volume is a little unusual for the series in that it has lots of pictures, but that's only because this is homological algebra, so there are many commutative diagrams. Most of the volumes are just walls of text (though the formatting and the production tend to be very clear).
(I believe they actually wrote some historical remarks in some of the books which were collected in a separate volume - I don't see any historical material in the volume I'm looking at, however. The members were not unmindful of things like history: Dieudonne wrote an excellent history of algebraic and differential topology, and Andre Weil wrote a book on the history of numbers.)
The fact that it took a while for many of the volumes to be translated from French to English may have deterred some English readers (though mathematical French is not too hard to understand even if you don't know French [like me]).
On the whole, (in my opinion) the presentation is too relentlessly formal for most people to try learning a subject (as opposed to a small topic) by reading Bourbaki. They did produce a "definitive exposition" of the subjects they covered, in the sense that the results and proofs are there. It's just that most people would have a hard time learning any of the subjects by reading through the books.
Wow, that's interesting. I guess that's like a US company being called "MRE". We would view that like a veteran's owned and operated company. Interesting.
And all the products would be "MRE-Phone", "MRE-Pod", hehehe :)
This is one of the things that everyone gets the reference, but it won't be good to admit it publicly. This quote is known to almost everyone born in that area, and it's the first thing that come to mind when you hear the name.
I use to read about the power border agents have over foreigners and was amazed at how easily they can destroy me. The only reason this hasn't happened seems to be that they're mostly decent, professional people. And now that's gone.
It's luck of the draw with those guys. The job attracts people with nationalistic/fascist mindsets and if you look a certain way you can expect to be treated worse than others.
I wouldn’t describe them as nationalistic or fascist. There’s no need to bring in any sort of political view… the problem is when little people get a little power. It can get ugly.
It’s the same people. There’s always been an element of luck.
I travel a lot and I’ve interacted with a lot of border agents. I’d say luck is important for a lot of countries. I’ve had great experiences flying in US, Egypt, UK and Turkey I’ve also had terrible experiences with US, and Iceland. Most other places have been somewhere in between.
In academic publishing, there is an implicit agreement between the authors and the journal to roughly match the importance of the paper to the prestige of the journal. Since there is no universal standard on either the prestige of the journal or the importance of the paper, mismatches happen regularly, and rejection is the natural result. In fact, the only way to avoid rejections is to submit a paper to a journal of lower prestige than your estimate, which is clearly not what authors want to do.
It’s not an accident - if academics underestimated the quality of their own work or overestimated that of the journal, this would increase acceptance rates.
Authors start at an attainable stretch goal, hope for a quick rejection if that’s the outcome, and work their way down the list. That’s why rejection is inevitable.
reply