87 Commits

Author SHA1 Message Date
antirez
6a0b1b5b43 Over 80 chars comment trimmed in pfcountCommand(). 2014-12-02 17:03:20 +01:00
antirez
d34fade2da Remove warnings and improve integer sign correctness. 2014-08-26 10:41:02 +02:00
antirez
dd50295de6 PFSELFTEST: less false positives.
This is just a quickfix, for the nature of the test the right way to fix
it is to average the error of N runs, since otherwise it is always
possible to get a false positive with a bad run, or to minimize too much
this possibility we may end testing with too much "large" error ranges.
2014-07-23 11:45:08 +02:00
Mike Trinkala
73737c4b85 Correct the HyperLogLog stale cache flag to prevent unnecessary computations.
Set the MSB as documented.
2014-05-19 15:45:22 +02:00
antirez
82556c1687 Speedup hllRawSum() processing 8 bytes per iteration.
The internal HLL raw encoding used by PFCOUNT when merging multiple keys
is aligned to 8 bits (1 byte per register) so we can exploit this to
improve performances by processing multiple bytes per iteration.

In benchmarks the new code was several times faster with HLLs with many
registers set to zero, while no slowdown was observed with populated
HLLs.
2014-04-18 16:14:34 +02:00
antirez
5e928eff64 Speedup SUM(2^-reg[m]) in HyperLogLog computation.
When the register is set to zero, we need to add 2^-0 to E, which is 1,
but it is faster to just add 'ez' at the end, which is the number of
registers set to zero, a value we need to compute anyway.
2014-04-18 16:14:34 +02:00
antirez
00726b2e0e PFCOUNT support for multi-key union. 2014-04-18 16:14:34 +02:00
antirez
57b34a73c2 HyperLogLog low level merge extracted from PFMERGE. 2014-04-18 16:14:34 +02:00
antirez
a3848a7429 HyperLogLog invalid representation error code set to INVALIDOBJ. 2014-04-16 15:09:47 +02:00
antirez
030336d4ef PFDEBUG TODENSE added.
Converts HyperLogLogs from sparse to dense. Used for testing.
2014-04-16 15:09:47 +02:00
antirez
e4c2afb9c0 User-defined switch point between sparse-dense HLL encodings. 2014-04-16 15:09:47 +02:00
antirez
e19e078c77 PFSELFTEST improved with sparse encoding checks. 2014-04-16 15:09:47 +02:00
antirez
66f67208d3 PFDEBUG ENCODING added. 2014-04-16 15:09:47 +02:00
antirez
db929041b5 Set HLL_SPARSE_MAX to 3000.
After running a few benchmarks, 3000 looks like a reasonable value to
keep HLLs with a few thousand elements small while the CPU cost is
still not huge.

This covers all the cases where the dense representation would use N
orders of magnitude more space, like in the case of many HLLs with
carinality of a few tens or hundreds.

It is not impossible that in the future this gets user configurable,
however it is easy to pick an unreasoable value just looking at savings
in the space dimension without checking what happens in the time
dimension.
2014-04-16 15:09:47 +02:00
antirez
90b67e416a Error message for invalid HLL objects unified. 2014-04-16 15:09:47 +02:00
antirez
fdace9bc50 PFMERGE fixed to work with sparse encoding. 2014-04-16 15:09:47 +02:00
antirez
159926a170 Correctly replicate PFDEBUG GETREG.
Even if it is a debugging command, make sure that when it forces a
change in encoding, the command is propagated.
2014-04-16 15:09:47 +02:00
antirez
fc33779283 Added assertion in hllSparseAdd() when promotion to dense occurs.
If we converted to dense, a register must be updated in the dense
representation.
2014-04-16 15:09:47 +02:00
antirez
67e15a0ff7 hllSparseAdd(): speed optimization.
Mostly by reordering opcodes check conditional by frequency of opcodes
in larger sparse-encoded HLLs.
2014-04-16 15:09:47 +02:00
antirez
ed52bbe134 Detect corrupted sparse HLLs in hllSparseSum(). 2014-04-16 15:09:46 +02:00
antirez
fd137b0781 hllSparseAdd(): faster code removing conditional.
Bottleneck found profiling. Big run time improvement found when testing
after the change.
2014-04-16 15:09:46 +02:00
antirez
8b3d5a8eb0 Comment typo in hllSparseAdd(). first -> fits. 2014-04-16 15:09:46 +02:00
antirez
46f9a78d9c Merge adjacent VAL opcodes in hllSparseAdd().
As more values are added splitting ZERO or XZERO opcodes, try to merge
adjacent VAL opcodes if they have the same value.
2014-04-16 15:09:46 +02:00
antirez
9f12467c9c More robust HLL_SPARSE macros protecting 'p' with parens.
Now the macros will work with arguments such as "ptr+1".
2014-04-16 15:09:46 +02:00
antirez
10bc0aea8d hllSparseAdd() opcode seek stop condition fixed. 2014-04-16 15:09:46 +02:00
antirez
8be9394ec7 Fixed error message generation in PFDEBUG GETREG.
Bulk length for registers was emitted too early, so if there was a bug
the reply looked like a long array with just one element, blocking the
client as result.
2014-04-16 15:09:46 +02:00
antirez
8cc09325c8 Fixed memmove() count in hllSparseAdd(). 2014-04-16 15:09:46 +02:00
antirez
da084bdf84 hllSparseAdd(): more correct dense conversion conditional.
We want to promote if the total string size exceeds the resulting size
after the upgrade.
2014-04-16 15:09:46 +02:00
antirez
a7dbd28cff hllSparseToDense(): sanity check added.
The function checks if all the HLL_REGISTERS were processed during the
convertion from sparse to dense encoding, returning REDIS_OK or
REDIS_ERR to signal a corruption problem.

A bug in PFDEBUG GETREG was fixed: when the object is converted to the
dense representation we need to reassign the new pointer to the header
structure pointer.
2014-04-16 15:09:46 +02:00
antirez
cf500998e8 PFDEBUG DECODE added.
Provides a human readable description of the opcodes composing a
run-length encoded HLL (sparse encoding).
The command is only useful for debugging / development tasks.
2014-04-16 15:09:46 +02:00
antirez
2c4a1eccda PFDEBUG added, PFGETREG removed.
PFDEBUG will be the interface to do debugging tasks with a key
containing an HLL object.
2014-04-16 15:09:46 +02:00
antirez
5786846122 hllSparseToDense API changed to take ref to object.
The new API takes directly the object doing everything needed to
turn it into a dense representation, including setting the new
representation as object->ptr.
2014-04-16 15:09:46 +02:00
antirez
68fb3019a0 hllSparseAdd() sanity check for span != 0 added. 2014-04-16 15:09:46 +02:00
antirez
e7e6aa49d0 Fix hllSparseAdd() new sequence replacement when next is NULL.
sdsIncrLen() must be called anyway even if we are replacing the last
oppcode of the sparse representation.
2014-04-16 15:09:46 +02:00
antirez
b964380817 Fix seqlen computation in hllSparseAdd(). 2014-04-16 15:09:46 +02:00
antirez
1c6671ab90 Abstract hllSparseAdd() / hllDenseAdd() via hllAdd(). 2014-04-16 15:09:46 +02:00
antirez
e3f5a38695 hllSparseSum(): multiply 1 * runlen for zero entries. 2014-04-16 15:09:46 +02:00
antirez
5b69b984b3 Macro HLL_SPARSE_XZERO_LEN fixed. 2014-04-16 15:09:46 +02:00
antirez
2c1fa11124 Fix HLL sparse object creation #2.
Two vars initialized to wrong values in createHLLObject().
2014-04-16 15:09:46 +02:00
antirez
0b9f09c4e0 Increment pointer while iterating sparse HLL object. 2014-04-16 15:09:46 +02:00
antirez
e526d32635 Fix HLL sparse object creation.
The function didn't considered the fact that each XZERO opcode is
two bytes.
2014-04-16 15:09:46 +02:00
antirez
ccbabf01e1 Create HyperLogLog objects with sparse encoding. 2014-04-16 15:09:46 +02:00
antirez
df5782f4d0 HyperLogLog sparse to dense conversion function. 2014-04-16 15:09:46 +02:00
antirez
e872394176 HyperLogLog sparse representation initial implementation.
Code never tested, but the basic layout is shaped in this commit.
Also missing:

1) Sparse -> Dense conversion function.
2) New HLL object creation using the sparse representation.
3) Implementation of PFMERGE for the sparse representation.
2014-04-16 15:09:46 +02:00
antirez
4e12ecdba5 hllCount() refactored to support multiple representations. 2014-04-16 15:09:46 +02:00
antirez
05b5032302 hllAdd() refactored into two functions.
Also dense representation access macro renamed accordingly.
2014-04-16 15:09:46 +02:00
antirez
4407ca6fc8 HyperLogLog refactoring to support different encodings.
Metadata are now placed at the start of the representation as an header.
There is a proper structure to access the representation.
Still work to do in order to truly abstract the implementation from the
representation, commands still work assuming dense representation.
2014-04-16 15:09:46 +02:00
antirez
099ae637ab HyperLogLog sparse representation slightly modified.
After running a few simulations with different alternative encodings,
it was found that the VAL opcode performs better using 5 bits for the
value and 2 bits for the run length, at least for cardinalities in the
range of interest.
2014-04-16 15:09:46 +02:00
antirez
2b531f729d HyperLogLog sparse representation description and macros. 2014-04-16 15:09:46 +02:00
antirez
f5bdaf366e PFCOUNT: always unshare/decode the object.
This will be a non-op most of the times since the object will be
unshared / decoded, however it is more technically correct to start this
way since the object may be decoded even in the read-only code path.
2014-04-16 15:09:46 +02:00