Binal Weerasena, reinforcement learning researcher from Sri Lanka

Reinforcement learning researcher · Sri Lanka

Binal builds AI that learns to make decisions under uncertainty

I'm Binal Weerasena, a Sri Lankan electrical engineer and AI researcher. I work on reinforcement learning, deep learning and data-driven decision-making, with applications in finance, energy and engineering systems.

Email me

Research

Reinforcement learning and AI research, from my MSc dissertation to studies I run in my own time.

Replay of the grid world: eight agents, items and three places, with what each agent said on turn 1
Independent research · 2026
LLM agentsMulti-agentPre-registered

Independent research · pilot with Claude Opus 5

Alpha Swarm: does a leader emerge among LLM agents?

Eight Claude agents share a 14 × 14 grid world for 40 turns. Three are given delivery objectives and five are told they have none; no prompt mentions leading, following or cooperating. The pre-registered study asks whether one agent comes to lead the others, and whether that leader is goal-driven. A blind LLM judge labels every request and reads each agent's private thoughts to count who acted on whose request. In the pilot, the agents delivered 17 of 18 items, 11 of them by agents with no objective, and no alpha emerged: the most-followed agent received 18% of follow events against a 30% threshold. Because goal-less agents took up the work unprompted, a minimal-prompt condition was added before the main runs, which now continue on local open models.

Side-by-side panels of what Q-learning, REINFORCE and actor-critic each learn on a 5 by 5 grid world
Independent research · 2026
Q-learningREINFORCEActor-critic

Independent study · learning project

Three families of reinforcement learning on one slippery maze

A hands-on comparison of the three families of reinforcement learning: value-based (Q-learning), policy-based (REINFORCE) and actor-critic, all solving the same 5 × 5 slippery grid world where a move succeeds only 80% of the time. Every agent is tabular, so what each one knows can be drawn directly: value heatmaps for Q-learning and actor-critic, and bare policy arrows for REINFORCE. Over 15 seeds, Q-learning reached the best greedy policy (+0.72), REINFORCE varied widely between seeds (standard deviation 0.76), and actor-critic learned fastest (564 episodes to threshold) and most consistently (standard deviation 0.014), which is why most modern deep RL is built on actor-critic methods.

Abstract temperature waves and a population of candidate solutions
BSc research · 2021–2022
OptimisationGenetic algorithmSimulink

Final-year design project · University of Moratuwa

Energy-optimal operation of chiller plants with genetic algorithms

HVAC chiller plants are among the largest energy consumers in commercial buildings. We framed plant operation as an optimisation problem and developed a genetic-algorithm method that selects chilled- and condenser-water set-points for minimum energy use. A Simulink model of the plant was used to validate the method numerically, showing significant energy savings over rule-based baselines.

Publications

  1. Illustration of a manuscript page with a portfolio drawdown chart

    In preparation

    Weerasena, K. B., & Priyadarshana, Y. H. P. P. (2026).

    Dynamic risk-free benchmarking for robust drawdown control in reinforcement-learning portfolio management: Evidence from six emerging and frontier markets, with a crisis-similarity criterion for cross-market transfer.

    Manuscript in preparation for submission to the International Review of Economics and Finance (Elsevier). Lead author.

Ventures

Startups I'm building or part of.

Experience

Industry roles in financial market technology and chip design software.

Sep 2024 – Dec 2025

Business Analyst and Product Owner

LSEG Technology (London Stock Exchange Group) · Sri Lanka

  • Contributed to the LTS trade-surveillance platform, authoring alert requirement specifications.
  • Supported the Central Counterparty (CCP) clearing solution for A5X, a new Brazilian exchange.
lseg.com ↗

May 2022 – May 2023

Application Engineer

Synopsys Inc., Global Customer Support · Sri Lanka

  • Supported industry customers on Electronic Design Automation (EDA) place-and-route tools, diagnosing and resolving complex technical issues under tight service schedules.
synopsys.com ↗

Education

Electrical engineering, then big data analytics and a research dissertation in reinforcement learning.

2024 – 2026

MSc in Big Data Analytics

Robert Gordon University, Aberdeen, UK

Dissertation awarded Grade A (Excellent). A grades in Big Data Programming, Data Analysis, Research Methods and the MSc Project.

rgu.ac.uk ↗

2017 – 2022

BSc Engineering (Hons) in Electrical Engineering

University of Moratuwa, Sri Lanka

Second Class (Upper Division) Honours, GPA 3.49 / 4.00. Minor in Entrepreneurship.

uom.lk ↗

Skills

Machine learning

Reinforcement learningPPO / PPO-LSTMMulti-task learningTransfer learningDeep learningExplainable AI (SHAP)Experimental design

AI and language models

LLM agent prototypingConversational agentsGenerative AI tooling

Programming

PythonPyTorchStable-Baselines3Gymnasiumscikit-learnpandasStreamlitGitLaTeX

Data

Financial time seriesWavelet featuresData warehousingCloud computing

Domains

Quantitative financePortfolio optimisationEDAPower and control systems

Let's talk research.

Collaborations, questions and good paper recommendations are all welcome.

binalweerasena.k@gmail.com