Applied Mathematics & Computer Science · Universidad del Rosario · Bogotá, Colombia
Project guide · Updated CV: coming · Email
I build simulations and learning experiments, then test what their comparisons actually establish: strong baselines, held-out evaluation and explicit controls.
My research question is whether automated judges and monitors measure what they claim to measure. Public work so far is on reinforcement learning, recurrent networks and human-coordination data. Next, I want to apply the same measurement tests to LLM judges and monitors in post-training and multi-agent systems. That last sentence is a plan, not a result.
Status: expected graduation 2028 (date to confirm). No publications yet; one manuscript is in preparation and has not been submitted.
| Project | What I did |
|---|---|
| SCRES reinforcement learning · simulation |
Supervised project: supply-chain simulation and Gymnasium/PPO experiments judged against a same-contract static comparator. The learned policy was ahead in only 2 of 10 checkpoint means and 2 of 60 held-out streams; the earlier advantage claim was withdrawn. Start with the same-contract verdict. |
| Motor-RNN connectivity computational neuroscience |
Neuromatch team project, then my independent equal-plasticity control, which holds the number of trainable recurrent edges equal across densities. Under that control the density ordering from the team's primary run reversed: the error of the sparsest network minus the mean of the denser ones changed sign in 8 of 8 seeds (95% bootstrap interval −0.113 to −0.068; exploratory, one architecture). Start with the control note. |
| Spectral analysis of cognitive labor human coordination |
My reanalysis of Andrade-Lotero and Goldstone's human-search experiment (the data belong to the original authors), with a past-only temporal audit. Predictive advantage is not established. Start with the temporal audit. |
- SCRES: the PPO advantage over a strong same-contract static policy did not survive the comparison, so the earlier claim was withdrawn.
- Spectral: early-prediction AUCs of 0.804 and 0.860 used overlapping time windows and are not validated. No predictive or spectral advantage is established.
- Motor-RNN: the primary density ordering reversed under equal trainable edges (8 of 8 seeds; exploratory, one architecture; see the control note).
- HeliOS: a collaborative educational RISC-V kernel in C. I am a contributor, not the sole author. Architecture notes and QEMU smoke tests; smoke tests are not formal verification.
- ContratIA Abierta: open-procurement data pipelines, FastAPI services and traceable signals for human review, not automated allegations. Human validation is still pending.
- ChaosLab: double-pendulum dynamics with numerical checks and an interactive presentation. A physics course project, not an ML result.
- Merged pull request: ForagersEnv.step() simulation pipeline, in another group's repository (merged 2026-02-28).
- A short note on judge calibration.
- A two-page reanalysis note on the spectral experiment.
Described in the CV and not published: a retrieval-augmented support prototype with heuristic answer/escalation routing (synthetic data); numerical-analysis coursework; and a computational audit of a marmot social-network model, done with a collaborator whose name is withheld until they agree.
Read the question and the stated limitations, inspect a saved result and its method, then follow the documented test path. The existing evidence can be inspected without refitting any model.
Tools: Python · PyTorch · NumPy/SciPy · NetworkX · SimPy · Gymnasium · C · SQL · FastAPI · Git · pytest


