Why We Moved Some Jobs from Sidekiq to RabbitMQ/Sneakers

Why We Moved Some Jobs from Sidekiq to RabbitMQ/Sneakers


Sidekiq is great and we still use it in most of our projects. This is not a “Sidekiq is bad” post. We moved specific projects to RabbitMQ with Sneakers because for those workloads it was a better fit. Here is why, and how we did the migration without stopping anything.

The problems we had

1. Redis memory

Sidekiq keeps its queues in Redis, and Redis is in-memory. Most of the time this is fine. The problem was spikes. A bulk push from a third party, a batch of ODS updates, or a few batch jobs queued at the same time, and Redis memory fills up fast.

We could have increased Redis memory, but paying for a bigger Redis all month to survive an occasional spike didn’t make sense. RabbitMQ writes queued messages to disk, so a few lakh jobs sitting in a queue is not a memory problem.

2. Memory leaks in workers

A Sidekiq process runs forever and all jobs run as threads inside it. If any job has a memory leak, the process keeps growing until the box runs out of memory, and it takes every other job on that process down with it.

With Sneakers each queue gets its own worker process. A leak in one worker stays in that worker, and we can restart just that one. This mattered more than it sounds. People with different levels of experience keep joining and leaving the team, and you cannot assume every piece of code is memory-safe forever. Hunting a leak in the middle of a sprint is expensive, so we wanted an architecture where a leak is annoying instead of an outage.

3. Forcing a small job payload

Sidekiq gives you object.delay.some_method. It is convenient, and Sidekiq docs tell you not to overuse it, but in practice developers will pass ActiveRecord objects as job arguments. Those get serialized with their internal state and associations, and the payload becomes big and slow for no reason.

RabbitMQ messages are just JSON. There is no way to accidentally push an AR object. You pass an id and look it up in the worker. The tool itself forces the cleaner pattern, which works better than code review reminders.

(This was Sidekiq 5/6 era. Newer Sidekiq has Sidekiq.strict_args! and Sidekiq 7 removed .delay altogether, so this point matters less today.)

4. Headroom for spikes

Sidekiq is faster per job, no argument there. But when the queue depth jumps by 10x we would rather have the jobs sitting safely on disk than provision Redis for the worst day of the year.

This one has a side effect which bit us, more on that below in alerting.

How we migrated

We did not rewrite jobs. We wrote wrappers so existing code kept calling the same methods, and the wrapper decided whether the job goes to Sidekiq or RabbitMQ.

Setup

# Gemfile
gem 'bunny'
gem 'sneakers'
# config/initializers/sneakers.rb
Sneakers.configure(
  amqp: ENV.fetch('RABBITMQ_URL'),
  exchange: 'app',
  exchange_type: :direct,
  durable: true,
  ack: true,
  prefetch: 10,
  threads: 10,
  workers: 2,
  log: Rails.root.join('log', 'sneakers.log')
)
Sneakers.logger.level = Logger::INFO

Publisher

One place that knows how to talk to RabbitMQ. Everything else goes through this.

# app/lib/rabbit_publisher.rb
class RabbitPublisher
  def self.connection
    @connection ||= Bunny.new(ENV.fetch('RABBITMQ_URL')).tap(&:start)
  end

  def self.publish(queue, payload)
    channel  = connection.create_channel
    exchange = channel.direct('app', durable: true)
    channel.queue(queue, durable: true).bind(exchange, routing_key: queue)
    exchange.publish(payload.to_json, routing_key: queue, persistent: true)
  ensure
    channel&.close
  end
end

Replacing .delay

This was the main one. Instead of user.delay.send_welcome_email, the job carries only class name, method name and id. The worker loads the record and calls the method.

# app/lib/async_dispatch.rb
module AsyncDispatch
  def async(method_name, *args, queue: 'default')
    RabbitPublisher.publish(queue, {
      class:  self.class.name,
      id:     id,
      method: method_name,
      args:   args
    })
  end
end

# app/models/application_record.rb
class ApplicationRecord < ActiveRecord::Base
  self.abstract_class = true
  include AsyncDispatch
end

Usage changes from this:

user.delay.send_welcome_email

to this:

user.async(:send_welcome_email)

And the worker on the other side:

# app/workers/method_dispatch_worker.rb
class MethodDispatchWorker
  include Sneakers::Worker
  from_queue 'default'

  def work(msg)
    payload = JSON.parse(msg)
    record  = payload['class'].constantize.find_by(id: payload['id'])

    if record.nil?
      Sneakers.logger.warn("skip: #{payload['class']}##{payload['id']} not found")
      return ack!
    end

    record.public_send(payload['method'], *payload['args'])
    ack!
  rescue => e
    Sneakers.logger.error("fail: #{payload.inspect} => #{e.message}")
    reject!
  end
end

ack! only after the method finishes, so if the worker dies mid-job the message goes back to the queue. If the record is gone by the time the job runs, we ack and move on instead of retrying forever.

Existing Sidekiq workers

For classes already written as Sidekiq workers we didn’t want to touch the perform method. We added a module that overrides perform_async to publish to RabbitMQ, and one generic Sneakers worker that calls perform.

# app/lib/rabbit_worker.rb
module RabbitWorker
  def self.included(base)
    base.extend(ClassMethods)
  end

  module ClassMethods
    def perform_async(*args)
      RabbitPublisher.publish(rabbit_queue, { class: name, args: args })
    end

    def rabbit_queue
      @rabbit_queue || 'workers'
    end

    def rabbit_options(queue:)
      @rabbit_queue = queue
    end
  end
end
# app/workers/sidekiq_bridge_worker.rb
class SidekiqBridgeWorker
  include Sneakers::Worker
  from_queue 'workers'

  def work(msg)
    payload = JSON.parse(msg)
    payload['class'].constantize.new.perform(*payload['args'])
    ack!
  rescue => e
    Sneakers.logger.error("fail: #{payload.inspect} => #{e.message}")
    reject!
  end
end

Migrating a worker was then one line:

class PriceSyncWorker
  include Sidekiq::Worker
  include RabbitWorker            # add this
  rabbit_options queue: 'pricing'

  def perform(property_id)
    # unchanged
  end
end

Because RabbitWorker is included after Sidekiq::Worker, its perform_async wins. Callers keep calling PriceSyncWorker.perform_async(id) and don’t know anything changed.

Running both in parallel

We did not flip everything in one day. Sidekiq kept running, and we moved job classes one at a time by including the module. Anything already sitting in Redis still got processed by Sidekiq, anything new went to RabbitMQ. After a few days of watching both, when the Sidekiq queues for those classes were empty and stayed empty, we removed those workers from the Sidekiq config.

Running the Sneakers side:

WORKERS=MethodDispatchWorker,SidekiqBridgeWorker bundle exec rake sneakers:run

For heavy queues we ran the worker for that queue separately so it could be scaled and restarted on its own:

WORKERS=PricingWorker bundle exec rake sneakers:run

Alert on queue depth

With Redis, a stuck job announces itself. Memory fills up, Redis starts failing, alerts fire, and you know within minutes. With RabbitMQ nothing fails. Jobs pile up on disk quietly, the machine is fine, the app is fine, and you find out a worker was stuck when someone asks why their email from yesterday never arrived. The same thing that gives you headroom also hides the problem.

So the first thing we added after migration was an alert on message count per queue. Simple cron, passive declare so it doesn’t create queues by mistake:

# lib/tasks/rabbit_check.rake
namespace :rabbit do
  desc 'Alert if any queue has too many messages'
  task check_depth: :environment do
    thresholds = { 'default' => 5_000, 'pricing' => 20_000, 'workers' => 5_000 }

    channel = RabbitPublisher.connection.create_channel
    thresholds.each do |queue, limit|
      count = channel.queue(queue, durable: true, passive: true).message_count
      Rails.logger.info("rabbit queue=#{queue} messages=#{count}")
      next if count < limit

      Alerter.notify("RabbitMQ queue #{queue} has #{count} messages (limit #{limit})")
    end
    channel.close
  end
end
*/5 * * * * cd /app && bundle exec rake rabbit:check_depth

Later we moved this to a Prometheus exporter against the management API (/api/queues) and alerted on both messages_ready and consumer count being zero. Consumer count zero is the important one: a high message count can just be a spike, but a queue with messages and no consumer means the worker died or is stuck.

Things to watch

  • Queue depth alerts are not optional. See above. Alert on consumer count too, not just messages.
  • constantize on data coming from a queue is only fine because we are the only publisher. Don’t do this if anything external can write to the queue.
  • Sidekiq’s retry with backoff is nicer out of the box. With Sneakers you get reject! and a dead letter queue, and you set up retries yourself.
  • Sidekiq web UI is better than the RabbitMQ management UI for looking at individual jobs. We lived with it.
  • Sneakers workers are still long-running processes. The isolation helps, but you still need to restart leaky ones. We set a memory limit per worker and let the supervisor restart them.

Summary

  • Redis is memory, RabbitMQ is disk. For spiky queues that is the whole argument.
  • One process per queue limits the blast radius of a memory leak.
  • JSON-only payloads stop people from pushing AR objects into jobs.
  • Wrappers over .delay and perform_async meant callers didn’t change.
  • Run both for a few days, move one job class at a time.
  • Disk-backed queues fail silently. Alert on queue depth and consumer count from day one.

Sidekiq is still the default for new projects here. This was for the ones where queue depth and memory were the problem, not throughput.