3696 Commits

Author SHA1 Message Date
antirez
d428569d13 More robust HLL_SPARSE macros protecting 'p' with parens.
Now the macros will work with arguments such as "ptr+1".
2014-04-16 15:26:27 +02:00
antirez
2d1e35a58a hllSparseAdd() opcode seek stop condition fixed. 2014-04-16 15:26:27 +02:00
antirez
29c7d1974b Fixed error message generation in PFDEBUG GETREG.
Bulk length for registers was emitted too early, so if there was a bug
the reply looked like a long array with just one element, blocking the
client as result.
2014-04-16 15:26:27 +02:00
antirez
023f2f7a10 Fixed memmove() count in hllSparseAdd(). 2014-04-16 15:26:27 +02:00
antirez
91391760ca hllSparseAdd(): more correct dense conversion conditional.
We want to promote if the total string size exceeds the resulting size
after the upgrade.
2014-04-16 15:26:27 +02:00
antirez
5d87160f8a hllSparseToDense(): sanity check added.
The function checks if all the HLL_REGISTERS were processed during the
convertion from sparse to dense encoding, returning REDIS_OK or
REDIS_ERR to signal a corruption problem.

A bug in PFDEBUG GETREG was fixed: when the object is converted to the
dense representation we need to reassign the new pointer to the header
structure pointer.
2014-04-16 15:26:27 +02:00
antirez
a18e7afbe8 PFDEBUG DECODE added.
Provides a human readable description of the opcodes composing a
run-length encoded HLL (sparse encoding).
The command is only useful for debugging / development tasks.
2014-04-16 15:26:27 +02:00
antirez
5c86d9e7b0 PFDEBUG added, PFGETREG removed.
PFDEBUG will be the interface to do debugging tasks with a key
containing an HLL object.
2014-04-16 15:26:27 +02:00
antirez
0e0ddcc4fa hllSparseToDense API changed to take ref to object.
The new API takes directly the object doing everything needed to
turn it into a dense representation, including setting the new
representation as object->ptr.
2014-04-16 15:26:27 +02:00
antirez
533235f7e3 hllSparseAdd() sanity check for span != 0 added. 2014-04-16 15:26:27 +02:00
antirez
f12fb49bb2 Fix hllSparseAdd() new sequence replacement when next is NULL.
sdsIncrLen() must be called anyway even if we are replacing the last
oppcode of the sparse representation.
2014-04-16 15:26:27 +02:00
antirez
a3e2adc0ae Fix seqlen computation in hllSparseAdd(). 2014-04-16 15:26:27 +02:00
antirez
eaa73ef73a Abstract hllSparseAdd() / hllDenseAdd() via hllAdd(). 2014-04-16 15:26:27 +02:00
antirez
95de3eaea0 hllSparseSum(): multiply 1 * runlen for zero entries. 2014-04-16 15:26:27 +02:00
antirez
c277169ed7 Macro HLL_SPARSE_XZERO_LEN fixed. 2014-04-16 15:26:27 +02:00
antirez
b6ca71a40a Fix HLL sparse object creation #2.
Two vars initialized to wrong values in createHLLObject().
2014-04-16 15:26:27 +02:00
antirez
8b39fee4c8 Increment pointer while iterating sparse HLL object. 2014-04-16 15:26:27 +02:00
antirez
c1ddead779 Fix HLL sparse object creation.
The function didn't considered the fact that each XZERO opcode is
two bytes.
2014-04-16 15:26:27 +02:00
antirez
0a7ba6faef Create HyperLogLog objects with sparse encoding. 2014-04-16 15:26:27 +02:00
antirez
8e889728bf HyperLogLog sparse to dense conversion function. 2014-04-16 15:26:27 +02:00
antirez
aadd513160 HyperLogLog sparse representation initial implementation.
Code never tested, but the basic layout is shaped in this commit.
Also missing:

1) Sparse -> Dense conversion function.
2) New HLL object creation using the sparse representation.
3) Implementation of PFMERGE for the sparse representation.
2014-04-16 15:26:27 +02:00
antirez
e5e5facba1 hllCount() refactored to support multiple representations. 2014-04-16 15:26:27 +02:00
antirez
88847a216f hllAdd() refactored into two functions.
Also dense representation access macro renamed accordingly.
2014-04-16 15:26:27 +02:00
antirez
21c97e71cd HyperLogLog refactoring to support different encodings.
Metadata are now placed at the start of the representation as an header.
There is a proper structure to access the representation.
Still work to do in order to truly abstract the implementation from the
representation, commands still work assuming dense representation.
2014-04-16 15:26:27 +02:00
antirez
ed96e115d0 HyperLogLog sparse representation slightly modified.
After running a few simulations with different alternative encodings,
it was found that the VAL opcode performs better using 5 bits for the
value and 2 bits for the run length, at least for cardinalities in the
range of interest.
2014-04-16 15:26:27 +02:00
antirez
93c75f1bf4 HyperLogLog sparse representation description and macros. 2014-04-16 15:26:26 +02:00
antirez
2b1385b9ec Add casting to match printf format.
adjustOpenFilesLimit() and clusterUpdateSlotsWithConfig() that were
assuming uint64_t is the same as unsigned long long, which is true
probably for all the systems out there that we target, but still GCC
emitted a warning since technically they are two different types.
2014-04-16 15:26:22 +02:00
antirez
12d1d18b59 ZRANGEBYLEX and ZREVRANGEBYLEX implementation. 2014-04-16 15:26:09 +02:00
antirez
a10bfded15 PFCOUNT: always unshare/decode the object.
This will be a non-op most of the times since the object will be
unshared / decoded, however it is more technically correct to start this
way since the object may be decoded even in the read-only code path.
2014-04-16 15:26:09 +02:00
antirez
ee764b0f8d Changed HyperLogLog hash seed to a non-zero value.
Using a seed of zero has the side effect of having the empty string
hashing to what is a very special case in the context of HyperLogLog: a
very long run of zeroes.

This did not influenced the correctness of the result with 16k registers
because of the harmonic mean, but still it is inconvenient that a so
obvious value maps to a so special hash.

The seed 0xadc83b19 is used instead, which is the first 64 bits of the
SHA1 of the empty string.

Reference: issue #1657.
2014-04-16 15:26:09 +02:00
antirez
091f5677a0 Initial HyperLogLog tests. 2014-04-16 15:26:09 +02:00
antirez
a97675923d Return "WRONGTYPE" error on PF* type mismatch. 2014-04-16 15:26:09 +02:00
antirez
faa7f259cc Fix PFADD infinite loop.
We need to guarantee that the last bit is 1, otherwise an element may
hash to just zeroes with probability 1/(2^64) and trigger an infinite
loop.

See issue #1657.
2014-04-16 15:26:09 +02:00
antirez
8ae6887091 Make hll-gnuplot-graph.rb callable from cli. 2014-04-16 15:26:09 +02:00
antirez
bdd8701c60 Remove HyperLogLog type checking duplicated code. 2014-04-16 15:26:09 +02:00
antirez
b3baa51403 PFGETREG added for testing purposes.
The new command allows to get a dump of the registers stored
into an HyperLogLog data structure for testing / debugging purposes.
2014-04-16 15:26:09 +02:00
antirez
348d5ea246 PFCOUNT: unshare the object when cached cardinality is modified. 2014-04-16 15:26:09 +02:00
antirez
ba0dc0be2c PFSELFTEST improved to test the approximation error. 2014-04-16 15:26:09 +02:00
antirez
c2cd48718d hll-gnuplot-graph.rb improved with new filter.
The function to generate graphs is also more flexible as now includes
step and max value. The step of the samples generation function is no
longer limited to min step of 1000.
2014-04-16 15:26:09 +02:00
antirez
ff551fad24 HyperLogLog: added magic / version.
This will allow future changes like compressed representations.
Currently the magic is not checked for performance reasons but this may
change in the future, for example if we add new types encoded in strings
that may have the same size of HyperLogLogs.
2014-04-16 15:26:09 +02:00
Raymond Myers
e08100a9ba Fixed pfadd/pfcount commands emitting hll* events instead of pf* events 2014-04-16 15:26:09 +02:00
Raymond Myers
23acae6e1c Change HLL* to PF* in error messages 2014-04-16 15:26:09 +02:00
antirez
e3773b0a0e Include redis.h before other stuff in hyperloglog.c.
Otherwise fmacros.h is included later and this may break compilation on
different systems.
2014-04-16 15:26:09 +02:00
antirez
badf23f57b HyperLogLog API prefix modified from "P" to "PF".
Using both the initials of Philippe Flajolet instead of just "P".
2014-04-16 15:26:03 +02:00
antirez
6e71459c73 Makefile.dep updated with hyperloglog.o deps. 2014-04-16 15:25:28 +02:00
antirez
2ee2bec2d6 HyperLogLog: make API use the P prefix in honor of Philippe Flajolet. 2014-04-16 15:24:15 +02:00
antirez
1e90c0066c HLLMERGE fixed by adding a... missing loop! 2014-04-16 15:23:21 +02:00
antirez
be7eef911b hll-gnuplot-graph.rb: added new filter "all". 2014-04-16 15:23:21 +02:00
antirez
f2e707b679 HyperLogLog apply bias correction using a polynomial.
Better results can be achieved by compensating for the bias of the raw
approximation just after 2.5m (when LINEARCOUNTING is no longer used) by
using a polynomial that approximates the bias at a given cardinality.

The curve used was found using this web page:

    http://www.xuru.org/rt/PR.asp

That performs polynomial regression given a set of values.
2014-04-16 15:23:21 +02:00
antirez
be7fe2b92b HLLMERGE implemented.
Merge N HLL data structures by selecting the max value for every
M[i] register among the set of HLLs.
2014-04-16 15:23:21 +02:00