Emir Ceylan

Machine learning — Medical AI — CS '27

Clinical risk models · Explainable AI · Bioinformatics · Full-stack web · Sabancı University '27 · LLM & RAG systems · Painter

01 / ABOUT

ABOUT

SYSTEM PROFILE

CLASS / 2027

NODE / ISTANBUL, TR

2027
EXPECTED GRAD
07+
ROLES & RESEARCH PROJECTS
ML
FOR HEALTHCARE

PROFILE_NARRATIVE.txt

I'm a Computer Science student at Sabancı University, class of 2027, and I build machine learning models for medicine.

Most of my work so far is clinical risk prediction: take a large electronic health record dataset, engineer features a clinician would recognise, train a calibrated gradient-boosted model, and prove it beats the scores hospitals already use. I care about calibration and explainability as much as AUROC, because a model nobody trusts doesn't get used.

Around that: undergraduate research on explainable deep learning for healthcare and on machine learning for biomedical alloys, a teaching assistantship in data science, and a couple of full-stack products built in teams. Right now I'm moving toward LLM and RAG systems, and building one to learn it properly.

I also paint. It isn't a metaphor for anything. I just like it.

02 /

WORK

/07

WHAT I ACTUALLY DO

Machine learning for healthcare: calibrated risk models on real clinical data, explainability that a physician can read, and the data engineering underneath.

Concretely: I turn raw hospital and genomic data into features, train and validate models with leakage-safe cross-validation, calibrate them, explain them with SHAP, and report them to clinical standards. I've also shipped full-stack products with a team, and I'm now building retrieval-augmented LLM systems.

  • Clinical ML
  • Explainable AI
  • Bioinformatics
  • Full-stack
  • LLM / RAG
  • Clinical risk model

    0.898AUROC

    Calibrated 30-day post-discharge mortality model for patients over 70, trained on 172k hospital admissions. Beat the standard clinical scores by +0.15 and +0.07 AUROC under grouped cross-validation.

    XGBoost · GroupKFold · SHAP · TRIPOD+AI

  • EHR feature pipeline

    77clinical features

    Engineered from a 28 GB raw electronic-health-record release. Every feature has a clinical justification, and the whole pipeline runs on a laptop.

    DuckDB · polars · pandas

  • Genetic variant classifier

    98.8% peak accuracy

    Diagnosis support for a rare inherited bone disorder from 3,000+ sequences. Four algorithms compared; a novel feature representation based on collagen amino-acid structure did the heavy lifting.

    scikit-learn · Biopython · cross-validation

  • Legal-tech hackathon

    2ndplace · ₺20k seed

    B2B product that scores litigation outcomes and cites relevant precedent for uploaded case documents. Built in a weekend with a team of five.

    Python · NLP · product pitch

FULL HISTORY ON REQUEST

I keep the detailed list of roles, projects and metrics off the public site on purpose. If you're hiring or supervising, ask and I'll send the full page plus a CV.

03 /

SKILLS

/07

ML & data

  • XGBoost
  • scikit-learn
  • SHAP
  • pandas
  • polars
  • NumPy
  • DuckDB
  • SQL
  • Matplotlib
  • Biopython
  • Jupyter

Languages

  • Python
  • C++
  • JavaScript
  • SQL

Web & infra

  • React
  • Vite
  • Django REST Framework
  • PostgreSQL
  • Docker
  • Firebase / Supabase
  • Astro
  • Git / GitHub
  • Jira
  • Ubuntu / Linux

Concepts

  • Machine learning
  • Deep learning
  • Explainable AI
  • Model calibration
  • Full-stack web
  • REST APIs
  • Statistics
  • Data visualisation

Learning now

Deliberately moving from classical ML into LLM engineering. Building a RAG project to prove it.

  • LLM systems
  • Retrieval-augmented generation
  • Agents

ALSO

Creative

  • Oil painting
  • Watercolour
  • Drawing
  • Digital illustration
  • Photoshop
  • Premiere Pro

Spoken

  • English (C1, fluent)
  • Turkish (native)