The world is running out of hooks
Why social ads stop working once people recognise the hook, and what that means as AI swallows most of the content we see.
Read on LinkedIn ↗
Reinforcement learning researcher · Sri Lanka
I'm Binal Weerasena, a Sri Lankan electrical engineer and AI researcher. I work on reinforcement learning, deep learning and data-driven decision-making, with applications in finance, energy and engineering systems.
Email meReinforcement learning and AI research, from my MSc dissertation to studies I run in my own time.
Portfolio agents trained with reinforcement learning tend to chase returns and suffer deep drawdowns, and they rarely have enough data in smaller markets. This dissertation develops a multi-task recurrent agent (PPO-LSTM) with a dynamic, risk-adjusted reward that benchmarks against the risk-free rate to control drawdowns. It transfers knowledge from data-rich to data-scarce markets and proposes a crisis-similarity criterion that predicts when cross-market transfer helps. The approach is validated on a six-market panel of emerging and frontier markets through multi-seed robustness tests, ablations, a factorial transfer × architecture design and SHAP-based interpretability.
Eight Claude agents share a 14 × 14 grid world for 40 turns. Three are given delivery objectives and five are told they have none; no prompt mentions leading, following or cooperating. The pre-registered study asks whether one agent comes to lead the others, and whether that leader is goal-driven. A blind LLM judge labels every request and reads each agent's private thoughts to count who acted on whose request. In the pilot, the agents delivered 17 of 18 items, 11 of them by agents with no objective, and no alpha emerged: the most-followed agent received 18% of follow events against a 30% threshold. Because goal-less agents took up the work unprompted, a minimal-prompt condition was added before the main runs, which now continue on local open models.
A hands-on comparison of the three families of reinforcement learning: value-based (Q-learning), policy-based (REINFORCE) and actor-critic, all solving the same 5 × 5 slippery grid world where a move succeeds only 80% of the time. Every agent is tabular, so what each one knows can be drawn directly: value heatmaps for Q-learning and actor-critic, and bare policy arrows for REINFORCE. Over 15 seeds, Q-learning reached the best greedy policy (+0.72), REINFORCE varied widely between seeds (standard deviation 0.76), and actor-critic learned fastest (564 episodes to threshold) and most consistently (standard deviation 0.014), which is why most modern deep RL is built on actor-critic methods.
HVAC chiller plants are among the largest energy consumers in commercial buildings. We framed plant operation as an optimisation problem and developed a genetic-algorithm method that selects chilled- and condenser-water set-points for minimum energy use. A Simulink model of the plant was used to validate the method numerically, showing significant energy savings over rule-based baselines.
In preparation
Manuscript in preparation for submission to the International Review of Economics and Finance (Elsevier). Lead author.
Why social ads stop working once people recognise the hook, and what that means as AI swallows most of the content we see.
Read on LinkedIn ↗Once ε-greedy has found the best machine, it keeps spending ε of its pulls on exploration. What a smarter strategy does instead.
Read on LinkedIn ↗A row of slot machines is the cleanest way into the explore-or-exploit trade-off at the heart of reinforcement learning.
Read on LinkedIn ↗Notes from my research: using transfer learning to carry what an agent learns in the US market over to the Colombo Stock Exchange.
Read on LinkedIn ↗What the Ministry of Finance's investor call on 11 February 2026 told international lenders about the debt report and the economy.
Read on LinkedIn ↗The Colombo Stock Exchange's Central Counterparty system makes the market safer and paves the way for derivatives. Are we ready?
Read on LinkedIn ↗Industry roles in financial market technology and chip design software.
Sep 2024 – Dec 2025
LSEG Technology (London Stock Exchange Group) · Sri Lanka
May 2022 – May 2023
Synopsys Inc., Global Customer Support · Sri Lanka
Electrical engineering, then big data analytics and a research dissertation in reinforcement learning.
2024 – 2026
Robert Gordon University, Aberdeen, UK
Dissertation awarded Grade A (Excellent). A grades in Big Data Programming, Data Analysis, Research Methods and the MSc Project.
rgu.ac.uk ↗2017 – 2022
University of Moratuwa, Sri Lanka
Second Class (Upper Division) Honours, GPA 3.49 / 4.00. Minor in Entrepreneurship.
uom.lk ↗
Collaborations, questions and good paper recommendations are all welcome.
binalweerasena.k@gmail.com