Hello Medmarks team,
I am a physician from Turkey building MedFailBench, an open source clinical AI safety benchmark based on synthetic clinician authored cases.
Project:
https://github.com/goktugozkanmd/medical-ai-failure-atlas
Medmarks looks like a strong technical fit because it already organizes runnable medical LLM benchmark environments and evaluation configs.
I am looking for technical collaboration with medical AI benchmark teams. The goal is to add a small clinical safety boundary evaluation layer that can complement existing medical benchmark scores.
MedFailBench focuses on cases where a model response may sound clinically fluent but crosses a safety boundary: missed urgent escalation, unsafe remote advice, unsupported certainty, weak source support, or unsafe safety net wording.
No patient data is involved. This is not clinical advice, not clinical validation, and not a deployment claim.
Would your team be open to discussing a small MedFailBench task integration or shared pilot?
Best,
Goktug Ozkan
Hello Medmarks team,
I am a physician from Turkey building MedFailBench, an open source clinical AI safety benchmark based on synthetic clinician authored cases.
Project:
https://github.com/goktugozkanmd/medical-ai-failure-atlas
Medmarks looks like a strong technical fit because it already organizes runnable medical LLM benchmark environments and evaluation configs.
I am looking for technical collaboration with medical AI benchmark teams. The goal is to add a small clinical safety boundary evaluation layer that can complement existing medical benchmark scores.
MedFailBench focuses on cases where a model response may sound clinically fluent but crosses a safety boundary: missed urgent escalation, unsafe remote advice, unsupported certainty, weak source support, or unsafe safety net wording.
No patient data is involved. This is not clinical advice, not clinical validation, and not a deployment claim.
Would your team be open to discussing a small MedFailBench task integration or shared pilot?
Best,
Goktug Ozkan