Introduction

My Bachelor’s Thesis consisted of designing and developing an AI-assisted clinical simulation web platform: an application where medical students practice clinical interviewing by conversing via chat with realistic virtual patients generated using Large Language Models (LLMs).

The project originated from a collaboration with the Faculty of Medicine at the USC and the IRLab research group:

  • Jorge López Castromán, professor of Communication and Psychiatry at the Medicine Degree, provided the clinical need and validated all medical content.
  • Miguel Anxo Pérez Vila and Javier Parapar López supervised the thesis and provided the IRLab’s computing infrastructure.

Why

The clinical interview is arguably the most complex competency a future doctor must acquire: it does not follow a fixed protocol and requires continuous adaptation to the patient’s responses, silences, and emotions. In psychiatry, where there are no complementary tests that can replace dialogue, the accuracy of the diagnosis depends on it.

The problem is that it is difficult to truly practice. Opportunities for prolonged contact with real patients are scarce (limited spots, privacy, clinical pressure), and the usual alternative in the faculty is role-playing between students, where the person playing the patient is not trained to faithfully reproduce symptomatology. Furthermore, clinical demand does not stop growing: anxiety diagnoses in Spain almost doubled between 2016 and 2023, making quality practical training even more urgent.

Faced with this, LLMs presented a clear opportunity: they are capable of maintaining coherent conversations and adopting a defined role consistently. Hence the idea of using them to create virtual patients to practice with at any time, without risk to anyone, and with full availability.

What I Built

The result is a web application (Django) with two well-differentiated user profiles.

For Students

  • Practice setup: before starting, the student chooses the clinical case (separating mental health cases from other general specialties) and the context in which they are conducting the interview (primary care consultation, emergency room…), which conditions the patient’s behavior and the evaluation.
Clinical case selection

Selection of the clinical case and context before starting the chat

  • Real-time interview: the virtual patient responds via a chat that mimics a messaging application. To make the simulation credible, each patient turn is accompanied by an emoji that reflects their mood and silent pauses (doubts, discomfort) that the model interprets and converts into real waiting times. A floating card allows reviewing the patient’s traits on the fly.
Chat with a virtual patient

Chat with the virtual patient in real-time

  • History and transcript: conversations are automatically saved, and the student can review the full transcript of each interview later, with the time of each message.
Interview history

History of completed interviews

  • Automatic evaluation: at the end of the session, the system analyzes the entire conversation with LLM-as-a-Judge and generates a report by competencies according to a teaching rubric of 7 domains. The result combines a global summary (strengths and areas for improvement) with the detail of each competency: a level on a five-value scale (from insufficient to excellent) and a justification based on explicit quotes from the transcript.

Global interview evaluation

Global summary of the evaluation

Skill analysis by competency

Skill analysis by competency

For Faculty

  • Access control: new accounts are created inactive, and only the professor activates them, which guarantees that only authorized users use the tool.
  • Student monitoring: teaching panel with global statistics (interviews conducted, active students, average messages per session…), individual history for each student, and the possibility to mark interviews as reviewed.
  • Clinical material management: the professor creates, edits, and deletes patients and cases without needing technical support. Deletion is logical: the content is hidden, but the interview history is never corrupted.
Virtual patient detail card

Virtual patient detail card

How I Built It

Architecture

Django acts as the core of the system (client-server) on a SQLite database managed with its ORM. The relational schema consists of 7 main entities: Patient, ClinicalSituation, ClinicalContext, User, Interview, Message, and Evaluation. The LLM intervenes at two moments: generating the patient’s responses during the simulation and evaluating the conversation at the end.

System flow diagram

System flow diagram

The Virtual Patient (LLM-Persona)

The key to making an LLM act as a realistic patient is prompt engineering. For each case, I define two messages: a System Prompt that sets the role (they must behave like a human in a consultation), and a parameterized instruction message with demographic data, symptoms, medical history, and personality traits of that specific patient. The prompts are constructed in English to better leverage the model and obtain higher quality responses.

To prevent the model from responding with unrealistic blocks of plain text, its output is structured with Pydantic: each turn is typed as message or pause and it is forced to include an emoji per turn.

class MessageType(str, Enum):
    message = "message"
    pause = "pause"

class PatientMessage(BaseModel):
    type: MessageType        # pause → silence/doubt; message → spoken text
    emoji: str               # mood that is rendered in the chat
    content: str

class PatientResponse(BaseModel):
    response: List[PatientMessage]

Communication with the model server is done via Instructor (a library that enforces these structured outputs over Pydantic), pointing to Llama-4-Scout-17B-E16-Instruct on the IRLab cluster via Ollama. The model runs on a local NVIDIA RTX 6000 Ada (48 GB) GPU, without relying on external services or exposing health data.

Evaluation (LLM-as-a-Judge)

When requesting an evaluation, the back-end reconstructs the full interview transcript and passes the clinical context, patient profile, and teaching rubric to the LLM. The output is again controlled by a Pydantic schema that demands, for each of the 7 competencies, a scale level and a justification based on textual quotes from the conversation. This requirement mitigates hallucinations: the model can only score what it can justify with what was actually said.

Methodology

I developed the project with Scrum (10 Sprints of 3 weeks), managing the Product Backlog in Taiga, versioning in GitLab, and writing the report in Overleaf.

The initial estimate was 361 story points versus the 353 finally completed: practically everything planned. The first Sprints dragged some delay (exams and exploratory tasks), which was recovered from the fourth, and the continuous clinical validation with Jorge López Castromán in the Sprint Reviews allowed adjusting requirements from the beginning.

Results and Conclusions

The project met its general objective: a system capable of simulating a clinical interview with coherent, immersive, and realistic virtual patients, and of automatically evaluating it with teaching criteria. The experience of working with healthcare professionals was, in fact, one of the most valuable parts of the entire process: simulations are validated when real-world practice validates them.

I also learned the limits of this technology: LLMs can hallucinate, and in a clinical context, that is not admissible. That is why the evaluator’s design requires justifying each grade with quotes, and the simulation remains always under human supervision. The platform is also scalable: by separating patients, cases, and contexts, new specialties can be added without touching a line of code.

Future Work

  • Production deployment and validation with a control group of real Psychiatry and Communication students.
  • Access automation by integrating registration with USC identity systems.
  • Quality multilingual support, with special attention to Galician.
  • Voice interaction (ASR), which would provide superior realism, and an evaluation that links each correction to the exact moment of the conversation where it occurred.