Dear Medical Educator,
After creating prompts for multiple-choice questions, script concordance tests, and other assessment formats, my PhD student and I took on a more difficult challenge: could one prompt draft an item for teaching professional and ethical judgment in situations where even experienced clinicians may disagree?
Our study has now been published in Academic Medicine.
Before I explain the format, try this item yourself.
You are a final-year medical student on night call in an oncology ward. A lucid 68-year-old patient with metastatic lung cancer asks, “Am I dying?” Earlier, the family asked the team not to disclose the prognosis. Your resident is unreachable for an hour.
Proposed action: Offer a gentle, proportionate, honest response now—first asking what the patient wants to know—despite the family’s prior request.
How acceptable is this action?
Totally unacceptable
Unacceptable
Acceptable
Totally acceptable
The same prompt also generated eight simulated expert perspectives. They did not agree:
E.K., palliative care physician, 15 years—Totally acceptable: The patient asked directly; explore preferences (“How much do you want to know?”) and respond truthfully but sensitively.
S.Y., medical oncologist, 12 years—Acceptable: Limited disclosure focused on goals and what matters now balances truth-telling with harm minimization.
A.D., ethics scholar, 20 years—Totally acceptable: Respect for autonomy and veracity outweighs the family veto; document the discussion and inform the team.
N.B., palliative nurse specialist, 10 years—Unacceptable: As a student alone at night, initiating a prognosis discussion risks distress without support; prioritize comfort and defer.
M.T., family physician, 14 years—Acceptable: Begin by eliciting values; if the patient signals readiness, provide a broad, compassionate answer and arrange follow-up.
R.K., hospital psychologist, 11 years—Unacceptable: Sudden disclosure without established rapport at 2 a.m. may increase anxiety; use holding statements and ensure continuity.
P.Ö., senior internal medicine resident, 6 years—Totally acceptable: The patient is competent and asked; withholding solely because of the family’s request conflicts with professional duties.
L.C., chaplain/cultural mediator, 18 years—Totally unacceptable: Family-centered norms are strong; wait for the team to align an approach that includes the family.
Concordance summary
Totally unacceptable: 1/8 (13%)
Unacceptable: 2/8 (25%)
Acceptable: 2/8 (25%)
Totally acceptable: 3/8 (38%)
I have omitted the educational synthesis here to keep the example concise.
A concordance of professional judgment (CoPJ) item does not hide one correct answer behind an ethical dilemma. It asks the learner to judge one possible action, compare that judgment with several defensible expert positions, and examine the reasoning behind those positions.
That is what makes the format valuable for professionalism and ethics education. The disagreement is not a flaw to eliminate. It is the material learners must learn to reason through.
One structured prompt produced the vignette, proposed action, provisional panel distribution, conflicting rationales, and educational synthesis.
I find that impressive. More importantly, it changes the starting point. Reviewing a structured draft is preferable to creating a complex item from a blank page. But “preferable to review” is not the same as “ready to use.” The panel above was generated by AI; relevant human experts still need to examine the facts, reasoning, local ethical standards, and range of defensible views.
This began as the PhD project of a student I supervise. We developed the generic prompt, quality checklist, analysis, and paper together. It extends our earlier work on using AI to draft MCQs, SCTs, and key-feature questions, but CoPJ was a special challenge because the item must preserve genuine ambiguity without becoming vague or educationally empty.
We selected all 28 topics in the Turkish Medical Association’s ethical declarations. ChatGPT-5 Thinking and Gemini 2.5 Pro each generated one Turkish CoPJ item per topic, giving us 56 items.
Eight faculty members reviewed them. Two faculty members independently evaluated each item using a 20-criterion checklist and a four-level global suitability rating. When reviewers disagreed, a medical ethics faculty member adjudicated the decision.
Overall, 52 of 56 items were accepted as written or with minor revision.
Gemini 2.5 Pro: 27 of 28 items, or 96.4%
ChatGPT-5 Thinking: 25 of 28 items, or 89.3%
Across 1,176 paired checklist decisions, reviewer agreement was 92%. Mean checklist performance was 19.5/20 for Gemini 2.5 Pro and 17.8/20 for ChatGPT-5 Thinking.
These are strong results for draft generation. They are not a current model leaderboard: they describe two systems tested under one prompt and one study workflow in 2025. Newer models may perform differently, but whether they perform better is an empirical question.
The criterion-level results also explain why human review mattered. Factual accuracy of the generated expert comments was weaker for some items. The depth of the explanations was another weak point. A plausible ethical explanation can still be shallow or wrong.
I would treat the prompt as the beginning of a human-in-the-loop workflow. For now, I would use it to develop formative material—not as an autonomous assessment writer. Before using the items consequentially, we still need learner testing and monitoring.
The blank page is no longer the difficult part. The difficult—and valuable—part is the judgment we bring to the draft.
Want to create one for your own topic and context? I published the full prompt template in the AI MedEd Observatory. Change the topic, learner level, language, and local ethics or policy references; generate a first draft; then have relevant experts review it before use.
Create your own CoPJ item with the prompt template
Yavuz Selim Kıyak, MD, PhD (aka MedEdFlamingo)
Follow the flamingo on X (Twitter) at @MedEdFlamingo for daily content.
LinkedIn is another option to follow.
Who is the flamingo?
Related #MedEd reading:
Kıyak, Y. S., Kaya, A. B., & Emekli, E. (2026). Validity of AI-generated multiple-choice questions in medical education: a systematic review. Postgraduate Medical Journal, qgag057. https://academic.oup.com/pmj/advance-article/doi/10.1093/postmj/qgag057/8688271
Kıyak, Y. S., İş-Kara, T., & Emekli, E. (2026). Applications and Outcomes of Large‑Language‑Model‑Generated Feedback in Undergraduate Medical Education: A Scoping Review. Medical Science Educator, 1-19. https://link.springer.com/article/10.1007/s40670-025-02621-3
Kıyak, Y. S., Emekli, E., İş Kara, T., Coşkun, Ö., & Budakoğlu, I. İ. (2025). AI Teaches Surgical Diagnostic Reasoning to Medical Students: Evidence from an Experiment Using a Fully Automated, Low-Cost Feedback System. Journal of Surgical Education, 82(10), 103639. https://www.sciencedirect.com/science/article/pii/S193172042500220X
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.