Science Cast

Testing GPT-4 with Wolfram Alpha and Code Interpreter plug-ins on math and science problems

Ernest DavisFebruary 26, 2024 11:36am

Views (141)
Comments (0)

Export Citation

Voice is AI-generated

Connected to paperThis paper is a preprint and has not been certified by peer review

Testing GPT-4 with Wolfram Alpha and Code Interpreter plug-ins on math and science problems

arXivPDF

Authors

Ernest Davis, Scott Aaronson

Abstract

This report describes a test of the large language model GPT-4 with the Wolfram Alpha and the Code Interpreter plug-ins on 105 original problems in science and math, at the high school and college levels, carried out in June-August 2023. Our tests suggest that the plug-ins significantly enhance GPT's ability to solve these problems. Having said that, there are still often "interface" failures; that is, GPT often has trouble formulating problems in a way that elicits useful answers from the plug-ins. Fixing these interface failures seems like a central challenge in making GPT a reliable tool for college-level calculation problems.

TwitterandLinkedIn

0 comments

Add comment

Testing GPT-4 with Wolfram Alpha and Code Interpreter plug-ins on math and science problems

Testing GPT-4 with Wolfram Alpha and Code Interpreter plug-ins on math and science problems

AI-powered Paper ChatBeta

AI-powered Paper ChatBeta

0 comments