FloRL

Implicit Normalizing Flows + Reinforcement Learning

Generate Convert Improve

Install / Use

/learn @joeybose/FloRL

About this skill

Quality Score

0/100

README

Improving Exploration in SAC with Normalizing Flows Policies

This codebase was used to generate the results documented in the paper "Improving Exploration in Soft-Actor-Critic with Normalizing Flows Policies". Patrick Nadeem Ward*12, Ariella Smofsky*12, Avishek Joey Bose12. INNF Workshop ICML 2019.

* Equal contribution, 1 McGill University, 2 Mila
Correspondence to:
- Patrick Nadeem Ward <Github: NadeemWard, patrick.ward@mail.mcgill.ca>
- Ariella Smofsky <Github: asmoog, ariella.smofsky@mail.mcgill.ca>

Requirements

Run Experiments

Gaussian policy on Dense Gridworld environment with REINFORCE:

TODO

Gaussian policy on Sparse Gridworld environment with REINFORCE:

TODO

Gaussian policy on Dense Gridworld environment with reparametrization:

python main.py --namestr=G-S-DG-CG --make_cont_grid --batch_size=128 --replay_size=100000 --hidden_size=64 --num_steps=100000 --policy=Gaussian --smol --comet --dense_goals --silent

Gaussian policy on Sparse Gridworld environment with reparametrization:

python main.py --namestr=G-S-CG --make_cont_grid --batch_size=128 --replay_size=100000 --hidden_size=64 --num_steps=100000 --policy=Gaussian --smol --comet --silent

Normalizing Flow policy on Dense Gridworld environment:

TODO

Normalizing Flow policy on Sparse Gridworld environment:

TODO

To run an experiment with a different policy distribution, modify the --policy flag.

References

Implementation of SAC based on PyTorch SAC.

Related Skills

YC-Killer

2.7k

A library of enterprise-grade AI agents designed to democratize artificial intelligence and provide free, open-source alternatives to overvalued Y Combinator startups. If you are excited about democratizing AI access & AI agents, please star ⭐️ this repository and use the link in the readme to join our open source AI research team.

groundhog

398

Groundhog's primary purpose is to teach people how Cursor and all these other coding agents work under the hood. If you understand how these coding assistants work from first principles, then you can drive these tools harder (or perhaps make your own!).

last30days-skill

13.8k

AI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary

000-main-rules

Project Context - Name: Interactive Developer Portfolio - Stack: Next.js (App Router), TypeScript, React, Tailwind CSS, Three.js - Architecture: Component-driven UI with a strict separation of conce

joeybose

View profile

View on GitHub

GitHub Stars62

CategoryEducation

Updated2mo ago

Forks7

joeybose/FloRL

Languages

Python

Security Score

95/100

Audited on Jan 15, 2026

No findings