← All communities
Reddit community

r/reinforcementlearning

87.5K members. The best post we have captured here reached 795 upvotes, which works out to 9 per 1,000 members.

9peak upvotes per 1,000 members
87.5Kmembers
795best post
66most comments
100posts we track

Open r/reinforcementlearning on Reddit β†—

What actually lands here

The mix of post types among everything we have captured. Posting a link where almost nothing but images lands is wasted effort, and this is the fastest way to see it.

38% 29% 23%
Video 38% Image 29% Text 23% Link 7% Question 3%

The topics it labels

Flairs are the community's own categories, and the closest thing a subreddit has to telling you what it wants.

Robot 6 D 6 R 4 Multi 3 N, P 2 N 2 DL, M, I 1 N, MF 1 Showcase 1 DL, MF, Multi, D 1

What the ceiling looks like

The biggest posts we have captured here. Links go to the original on Reddit.

We beat Google Deepmind but got killed by a chinese lab

↑ 795 πŸ’¬ 33 text 2025

Trained a PPO agent to beat Lace in Hollow Knight: Silksong

↑ 552 πŸ’¬ 66 video 2025

IT'S LEARNING!

↑ 549 πŸ’¬ 23 image 2025

PPO Ping Pong

↑ 359 πŸ’¬ 25 video 2025

Why is RL fine-tuning on LLMs so easy and stable, compared to the RL we're all doing?

↑ 355 πŸ’¬ 42 question 2025

Andrew G. Barto and Richard S. Sutton named as recipients of the 2024 ACM A.M. Turing Award

↑ 348 πŸ’¬ 14 link 2025

Can RL redefine AI vision? My experiments with partial observation & Loss as a Reward

↑ 324 πŸ’¬ 60 video 2025

Built a custom robotic arm environment and trained an AI agent to control it

↑ 314 πŸ’¬ 18 video 2025

Looking to improve Sim2Real

↑ 305 πŸ’¬ 33 video 2025

OpenAI Gym is now actively maintained again (by me)! Here's my plan

↑ 291 πŸ’¬ 33 text 2021

Communities its size

Within four times the member count either way, so the comparison is fair. Sorted by upside.

See what is climbing here before you write

ViralHunt tracks Reddit and eight other networks, records how every post grows, and puts what is taking off in front of your team in one list.

Start free β†’

Every figure here is what ViralHunt has captured from r/reinforcementlearning, not everything the community has ever posted. We deliberately do not publish an average score: our sample is weighted toward a community's best posts, so an average would be flattering and wrong. Peak per 1,000 members is computed the same way for every community listed. See how we measure.