myofficehours.ai

Issue 7 · 2 October 2026

myofficehours.ai — Issue 7

02 October 2026 · A weekly read on AI for faculty

When a finished product costs nothing, the evidence has to be made in the room

This week the people who run courses and PhD programmes started moving proof of learning out of take-home work and into live, timed, unassisted settings. A Calgary professor graded live group work, a Harvard summit of 24 mathematicians drafted in-person assessment for the PhD, and a University of Washington chair proposed live qualifying exams. In the same days, a near-top AI model became free to any student with a browser.

The free-model fact is the pressure behind the rest. Simon Willison reports that Anthropic's Claude Sonnet 5.5 now powers the free tier of claude.ai, while OpenAI's free tier still runs a weaker model. Assumptions in your assignments and syllabus policy may already be out of date. OpenAI also says its new GPT-6.1 Sol nearly matches its top model at a fifth of the price. That is OpenAI's own testing, not an independent result. The responses worth copying are small and specific. They do not try to detect AI. They change where the evidence of thinking is produced.

What people are actually trying

Sarah Elaine Eaton rebuilt an online Master's course so that 75 percent of the grade came from three timed group tasks in mandatory Zoom sessions. Students read beforehand, applied the readings in groups for five to ten minutes, and uploaded a shared artifact to the course dropbox. She graded visible thinking, including unresolved disagreements, and allowed AI use. It is one instructor's account, and the student reactions are her summary of informal feedback. The Harvard summit's draft recommendations are open for comment at education-summit@cmsa.fas.harvard.edu. Jess Werk, who chairs Astronomy at the University of Washington, proposes live unassisted qualifying exams, rubrics scored independently by each committee member, and counting PhD mentoring in promotion. She cites a preprint finding that over half of 2025 astronomy papers were AI-assisted, with about one in 66 declaring it.

The counter-current

The best randomised evidence this week cuts against easy enthusiasm. A meta-analysis of AI surgical training found AI instruction statistically indistinguishable from good human teaching, with an effect of 0.13. It looked large only against passive or no-AI comparisons, and certainty was rated low. Avi Staiman, whose firm sells language editing, collects studies showing that AI detectors disproportionately flag second-language writing, and that ChatGPT edits can quietly change the strength of a scientific claim. A summary of a new paper says scientists report saving about seven hours a week, but also a growing backlog of untested hypotheses and demand for verifying AI output. The paper's publication status is not stated. Open-weight agents are still well behind the closed frontier: Hcompany's Holo4 scores 61.7 percent on a desktop-task test where the top closed model scores 81.8.

For administrators

OpenAI launched always-on 'dots' agents on 29 September. Education workspaces get a beta only when the administrator enables it, so that decision is yours to make deliberately. Confirm the no-training-by-default promise in your Edu contract before you do. Premium Google Colab compute is moving into Google AI subscription plans, so check what students are told to buy. Sydney's reworded graduate qualities are a consultation draft, not adopted policy.

One thing to try this week

Paste one of your take-home assignment prompts into the free tier of claude.ai and see what comes back. Then take the weakest part and rewrite it as a ten-minute timed pair task in class, ending in a shared artifact students hand in. Grade the thinking on the page, as Eaton does. It takes under thirty minutes.


What someone who studies this thinks

O'Malley reads four new peer-reviewed studies and draws one conclusion: AI teaching has so far beaten poor or absent teaching, not good teaching. So every deployment decision is really a question about what the tool replaces. Swapping a supervised session for a chatbot is not supported by this evidence. Filling the hours when students currently get nothing probably is. He treats the large passive-comparator result as exploratory, notes that trained Birmingham students trusted AI less rather than more, and flags that the ethics framework he reviews has never been validated.

— Andrew O'Malley, Senior Lecturer, School of Medicine, University of St Andrews · read the piece


Where the experts actually disagree

If AI can produce any take-home product, what should a department change: the assessments, the faculty reward system, the graduate-quality wording, or the policing?

Sarah Elaine Eaton (pragmatist) — Change the assessment. Move most of the grade to live, timed work in session, grade the thinking, and permit AI use rather than fight it. Their argument

Jess Werk (skeptic) — Student rules are not where the leverage is. Departments should protect doctoral training with live unassisted exams, no AI-written text between advisor and student, and promotion criteria that weight mentoring. Their argument

Adam Bridgeman (neutral) — Do not add a standalone AI-fluency quality. Reinterpret the existing graduate qualities around judgement and human agency, because AI has raised the standard of proof rather than changed what is valued. Their argument

Avi Staiman (pragmatist) — 'Did you use AI?' is the wrong question. Specify which uses need verification and disclosure, and stop treating polished English as evidence of misconduct. Their argument

The position with no support this week is the one many institutions quietly hold: that suspicious prose or a detector score is adequate evidence of who did the work. Staiman's cited studies show detectors flag second-language writers and reviewers misread 'AI-sounding' style. Eaton and Werk move the evidence into the room instead. The open dispute is whether the fix is course-level assessment or department-level incentives, and the evidence this week shows both are still untested.


Threads we have been following

Issue 5, 2026-09-10 — we said: Issue 5 closed on a question: given that AI can solve almost any written assignment and cannot be trusted to grade one, what is the right response in a university course?

This week produced the first concrete answers. Eaton moved 75 percent of a course grade into timed live group tasks, a Harvard summit of 24 mathematicians drafted in-person assessment for the PhD, and Werk proposed live unassisted qualifying exams. All are one instructor's account, a draft, or a proposal, and none is evidence that it works. The premise also hardened, because a near-frontier model is now free to students.

Issue 3, 2026-08-14 — we said: Issue 3 said detection had moved to the vendors and that they were not ready, and warned that making a tool available is not the same as it being used, and using it is not the same as it helping.

Staiman's essay, from an interested party but built on named studies, adds a cost to 'helping'. It cites a study finding AI detectors disproportionately flag second-language writing. It also cites a roughly 80,000-review analysis in which bias against authors from less English-dominant countries persisted, with reviewers shifting to complaints about 'AI-sounding' phrases.


The rest of the week, briefly

One line each, ordered by how much it should change what you do. The full account of any of them is on the dashboard.

Act on this

Watch

12 more stories ran this week and are waiting on the dashboard.


What is gaining ground

Topics appearing more often across the sources we track than they did in the weeks before. Counted, not guessed.

  • introducing — 6 mentions, new this period
  • systematic reviews — 4 mentions, 8.0x more often
  • work — 4 mentions, 2.7x more often
  • age — 3 mentions, 3.0x more often
  • frontier — 3 mentions, 6.0x more often

Three worth your time

20 resources went into the library this week. These three are the ones to open first.

Generative AI Tools for Literature Researching: Comparison of AI Literature Review Tools — ailit · 15 minutes · Oregon State University Libraries (LibGuide by librarian Adam Lindsley)
After working through this guide, readers can compare Consensus, Elicit, Semantic Scholar and Research Rabbit across access and fees, features, drawbacks, privacy practices, data sources and reference-manager integration, and choose the tool that fits their literature review workflow. The OSU librarians' star ratings and candid drawbacks, such as Consensus sometimes misreading sources and Semantic Scholar's potential factual errors and text hallucinations, help readers judge each tool's reliability rather than take vendor claims at face value. Readers also learn the real cost before starting, since Semantic Scholar and Research Rabbit are free while Consensus and Elicit depend on institutional subscriptions or freemium accounts for their fuller features.

How to Use ChatGPT Edu — Custom workflows · 15_minutes · University of Iowa Information Technology Services
Readers can follow step-by-step instructions for logging into their institution's ChatGPT Edu workspace through SSO, confirming they are in the university workspace rather than an unprotected personal account, and requesting a licence, which the page prices at $156 per user per year. They can create a Skill from a prompt they already reuse to standardise repeated tasks such as summarising research articles, drafting emails, or generating rubrics, and they can decide how broadly to share a Custom GPT while keeping its contents to University Public or University Internal data. They will also come away knowing that chats are deleted after six months without interaction, that voice and agent features are unavailable, and that image generation and deep research are capped daily, so they can plan which workflows and data belong in the tool.

UVA Library AI Challenge: Step 9: AI and Presentations — Slides and decks · about 20 minutes · University of Virginia Library
Readers can run a ready-made prompt in Copilot Chat (or a free-tier alternative such as Canva AI or Slides AI) to generate slide titles, bullet points, and speaker notes for a short talk, then use follow-up prompts to expand the topic, refine the notes, and suggest visuals. Reflection questions ask them to assess how well the tool structured their message, whether they would feel comfortable presenting co-authored material, and how to keep their own voice in the content. The page also cautions against sharing private data with unlicensed tools and lists further reading plus a 14-minute video on AI-assisted presentations.


Who we read this week

This issue drew on Terence Tao, OpenAI, Google DeepMind, Hugging Face, Sarah Elaine Eaton, Lisa Janicke Hinchliffe, Tyler Cowen, Danny Liu. The full watchlist, with what each source is good for and where they stand on AI in education, is on the dashboard under "By voice".


The whole library lives on the dashboard, sorted by what you are trying to get done and by the tools you already have. Each task runs from a twenty-minute start to something you could spend a weekend on.

myofficehours.ai is assembled automatically: a daily sweep for new tutorials and a weekly edition every Friday. Every link is checked before it ships. Reply with anything broken, missing, or worth adding.