This PR avoids wrapping the return value from the next request handler in another promise. The request handler will either return a promise or a value that can be consumed directly. When running example 99 with included benchmarking instructions, this simple change improves performance from ~2600 req/s to ~2700 req/s on my local machine. Also, running the (very synthetic) tests/benchmark-middleware-runner.php improved from ~2s to ~1.7s on my local machine.
There's potential for a BC break here, but given the lack of documentation for the old behavior and its surprising semantics (see #287), I do not consider this to be a BC break. Instead, this PR now adds documentation for consuming the response from the next middleware request handler function to avoid any future BC breaks.