Gradient as a teaching tool: Lessons from a classroom pilot
Sep 10, 2026 · 6 min read
MB SamuelFounder
Can students grow their AI fluency from three weeks of working with AI?
This summer, our team at Gradient partnered with Dr. Steven Strauss, a professor who has taught courses at Harvard and Princeton, and teaching assistant Dean Marchildon to find out. Over their three-week course, we used Gradient, our tool for measuring AI fluency, to capture how students worked with AI at the start, how they worked by the end, and the growth in between.
About the course
Dr. Strauss and Dean were teaching a three-week, graduate-level summer course, Innovating with Generative AI for Leaders and Managers. The focus was hands-on: giving students real practice using AI to work through business and leadership problems.
We connected with them after Dr. Strauss published a thoughtful piece on how he defines AI fluency. His principles line up closely with how Gradient thinks about AI fluency: that it is not just about a grasp of prompting, tools, and workflows. It also means keeping your own independent judgment, and reflecting on your work so you continue to improve.
The goal of our collaboration was to use Gradient as a teaching tool to help students see the value of judgment and reflection in their AI use, and to measure how they improved.
How we used Gradient
We ran Gradient as a bookend around the course. Students took a Gradient exercise in the first few days, spent three weeks learning, and then took a second exercise at the end. This gave us a clean before-and-after for each student: the first exercise was the baseline, the second showed what had changed.
- Jul 13Course beginsThree-week intensive starts
- Jul 16Pre-testFirst Gradient assessment
- Jul 16 to 28CourseworkTwo weeks of intensive learning
- Jul 29Post-testSecond assessment, new scenario
- Aug 4Results89% improved, all nine skills up
Both the pre- and post-tests followed the same format. After a few minutes to get set up, students worked through three timed phases: triage a crowded inbox with AI to find what mattered most, build a repeatable workflow to do it every week, and reflect on how it went.
- Explore5 minGet set up in the workspace.
- Triage15 minSort a crowded inbox for what matters most.
- Build15 minTurn the triage into a repeatable weekly workflow.
- Reflect10 minWrite a short reflection on the experience.
Afterward, Gradient's scoring engine evaluated two things for each student: their process (their prompts, their tool calls, and how closely they engaged with the emails themselves) and the finished product from each phase. From both, it produced personalized feedback on each student's work and their AI fluency.
For the post-test, we kept the format and the difficulty identical but changed the scenario and the specific challenges hidden inside it. That was deliberate. We wanted to measure whether students had grown their AI skills, not whether they remembered the first test.
The results
Looking at the pre- and post-test scores side by side, the results were clear: students improved a lot over the three-week course.
The headline: 89% of students scored higher on the post-test than on the pre-test, and every one of the nine skills we measured went up across the cohort.
Here is what that looked like for AI fluency:
- Setup+14
The strongest skill in July, but still grew. Students actively managed context and customized.
80%94%+14 - Direction+11
More correcting, more re-running, more comparing one run against the next.
78%89%+11 - Judgment+9
Students verified the AI's work more the second time around, but still have room to grow.
63%72%+9 - Mindset+7
Reflections got more candid and more specific about what students would change.
62%69%+7
The clearest gains were in the habits that separate careful AI use from careless AI use:
- Students gave clearer instructions and established their goals up front more consistently on the post-test.
- Students verified AI outputs and cited their sources more often.
- They engaged more deeply with the workspace itself. They were much more likely to build reusable skills, customize the agents.md, and shape the environment to fit the task instead of taking the AI's first answer.
- Built a reusable skillMore than half of students made or brought in skills that helped them customize how the AI worked.Pre27%Post52%
- Wrote or edited agents.mdStanding instructions for the AI, written once and applied to every later run.Pre33%Post48%
- Edited the starter workflowRewriting the inherited rules rather than working around them.Pre70%Post81%
- Edited a skill after testing itStudents started to show positive signs of verification, like iterative testing.Pre12%Post22%
What we learned from the students
We learned as much from the students as they learned from the course.
Dr. Strauss asked students to write one more reflection after they received their Gradient feedback reports, and shared those takeaways with us. Two themes came up again and again.
The first was the importance of verifying AI outputs. One student put it plainly:
"A good instruction is not enough unless I test whether the system actually follows it."
The second was a shift in how students related to the tool. They stopped treating AI like a search box and started treating it like a junior colleague:
"Going forward, I would manage AI more like a capable but inexperienced team member. I would give clearer instructions, confirm that the tools and data sources are actually working, review the intermediate outputs, and test automated workflows on a small sample before relying on them."
Here's what Dean Marchildon, the teaching assistant for Innovating with Generative AI for Leaders and Managers, shared about the experience:
"This gave our students a valuable opportunity to build and assess their AI fluency through hands-on practice. It's one thing to talk about AI fluency, but another to apply it, make your thinking visible, recognize where your approach fell short, and clearly articulate what you learned. The growth in students' skills and confidence from the first assessment to the second (over just three weeks) was clear, and it was incredibly valuable to have a platform that could capture and measure that progress."
Students also gave us sharp product feedback: which parts of the exercise were confusing, and where the sandbox got in their way. We used those observations to inform a new onboarding flow, designed to make the Gradient environment easier to step into.
What comes next
While this was a small sample size, and not a rigorous controlled trial, these results were encouraging: with a few focused weeks of AI learning, we saw that students were able to score meaningfully higher on Gradient's AI fluency assessments.
What will be most important, though, is how students take these skills and apply them to increase their impact in life and work after the course.
Working with us
We are grateful to Dr. Strauss and Dean Marchildon for the partnership, and excited to keep exploring how Gradient can serve as both a teaching tool and a way to measure the impact of AI learning.
If you run a course or lead a team and you want to measure a baseline and prove growth like this, we would love to talk. Please reach out.
Start hiring AI fluent talent.
We’ll show you an assessment, walk through the scoring engine, and get you live in a few hours.