sql – Grey Panthers Savannah https://grey-panther.net Just another WordPress site Mon, 06 Apr 2009 09:17:00 +0000 en-US hourly 1 https://wordpress.org/?v=7.1.2 206299117 NoCOUG SQL challenge https://grey-panther.net/2009/04/nocoug-sql-challenge.html https://grey-panther.net/2009/04/nocoug-sql-challenge.html#comments Mon, 06 Apr 2009 09:17:00 +0000 https://grey-panther.net/?p=326 2724771924_3bd95fe601_o NoCOUG (which stands for Northen California Oracle Users Group) published an SQL challenge [PDF]: using SQL determine the probability of achieving a given number by throwing a non-balanced dice N times.

Being a PostgreSQL fanboy that I am, I’ve given a try with PG. Here are the results:

To create the table and populate it (note the 1.0 notation, otherwise it does integer arithmetic):


CREATE TABLE die
(
  face_id integer NOT NULL,
  face_value integer NOT NULL,
  probability double precision NOT NULL,
  CONSTRAINT die_pkey PRIMARY KEY (face_id)
);

INSERT INTO die(face_id, face_value, probability) VALUES
	(1, 1, 1.0/6 + 1.0/12),
	(2, 3, 1.0/6 + 1.0/12),
	(3, 4, 1.0/6 + 1.0/12),
	(4, 5, 1.0/6 - 1.0/12),
	(5, 6, 1.0/6 - 1.0/12),
	(6, 8, 1.0/6 - 1.0/12);

And here is a stored procedure written in plpgsql to calculate the probabilities:


CREATE OR REPLACE FUNCTION calc_probs(depth integer) RETURNS SETOF die AS $$
BEGIN
  IF (depth <= 1) THEN
    RETURN QUERY SELECT * FROM die;
  ELSE
    RETURN QUERY
      SELECT 
          MIN(A.face_id) AS face_id,
          A.face_value + B.face_value AS face_value, 
          SUM(A.probability * B.probability) AS probability 
        FROM die A, (SELECT * FROM calc_probs(depth - 1)) B
        GROUP BY 2;
  END IF;
END;
$$ LANGUAGE plpgsql;
SELECT face_value, probability FROM calc_probs(100) ORDER BY face_value;

I didn’t do any benchmarks, but it should be quite fast. One optimization you could do for such functions in production is to declare it STABLE (or an unsafe optimization: declare it IMMUTABLE if the underlying table changes very infrequently). From the documentation:

STABLE indicates that the function cannot modify the database, and that within a single table scan it will consistently return the same result for the same argument values, but that its result could change across SQL statements. This is the appropriate selection for functions whose results depend on database lookups, parameter variables (such as the current time zone), etc.

Finally, here is a solution using the CTE (Common Table Expression) feature from the upcoming 8.4 release (you can think of the CTE’s like dynamically defined VIEWS – for more details about them you can start at the following links: CTEReadme, Common Table Expressions (WITH and WITH RECURSIVE) and Waiting for 8.4 – Common Table Expressions (WITH queries)):


WITH RECURSIVE p AS (
    SELECT 1 AS throws, face_value, probability FROM die
UNION ALL
    SELECT B.throws + 1 AS throws,
        A.face_value + B.face_value AS face_value,
        A.probability * B.probability AS probability
        FROM die A, p B
        WHERE B.throws < 3
) SELECT face_value, SUM(probability) FROM p
    WHERE throws = 2 GROUP BY face_value ORDER BY face_value;

This solution is a little less optimal because it does the GROUPing at the end, but I wasn’t able to include it in the inner select (it kept giving me an error saying that grouping is not supported on the recursive part of the queries). It also makes you repeat the number of dice throws twice. It is possible that it can be written better (this is the first time I’ve experimented with the particular feature) or that it just isn’t suitable for these types of queries.

Toying around with this challenge was fun and it certainly shows that PostgreSQL is on par with most of Oracle’s features.

Picture taken from TooFarNorth’s photostream with permission.

]]>
https://grey-panther.net/2009/04/nocoug-sql-challenge.html/feed 2 326
Build a botnet – without infecting end-users https://grey-panther.net/2009/03/build-a-botnet-without-infecting-end-users.html https://grey-panther.net/2009/03/build-a-botnet-without-infecting-end-users.html#respond Thu, 26 Mar 2009 14:30:00 +0000 https://grey-panther.net/?p=336 31219031_449e05f104_b The idea is not new: get a lot of users to view a given webpage, to DDoS the webserver / backend (depending where the bottlenecks are). If I recall correctly, some student asked the visitors of his website to continuously refresh the page of his university and got charged for it.

As many have remarked at the time (a) the university had some weak webservers if it caved to such simple methods and (b) this can be done automatically with Javascript or Flash and would be very hard to track down.

Imagine the following scenario:

  • The attacker inserts arbitrary Javascript or Flash content on one or more medium-to-high traffic websites. This can be done multiple ways: one can hack into CMS’s and modify the content of the articles to include the code in the articles. There are many vulnerable sites out there. Or, an even simpler solution is to buy placements for Flash banner ads and include the code in them.
  • The code (a) looks up a DNS name (this makes the attack targetable) (b) launches N “threads” and starts sending requests to the given website

Such attacks would be very hard to diagnose. The requests would come intermittently from a wide range of IP addresses. Even if you could get your hands on such a computer, you couldn’t find the source of the requests easily (it’s not like the computer is infected with a malware you can find by scanning the files on the hard-disk). It can be also very sneaky, randomly executing (or not) or using geotargeting to select a subset of computers. These techniques are already in use by malicious advertisements (“malwertisements”) which are currently used to try to sell you rogue AV products. An other reason which makes finding the source hard, is the fact that AFAIK XMLHttpRequest does not send the referrer header. An other way to get rid of the referrer header is to make the request from a HTTPS site (browsers do not send referrer in this situation to avoid information leakage).

What can you do? Not very much. Prepare for the DDoS. Have a contingency plan (like a backup location in a different IP space and pointing your DNS entry there). You might be able to differentiate the requests from “normal” requests, but even so, the volume of requests can bring down the machine at the TCP level. And please, please secure your website. We have enough unsecured websites already!

Picture taken from 416style’s photostream with permission.

]]>
https://grey-panther.net/2009/03/build-a-botnet-without-infecting-end-users.html/feed 0 336
Efficient SQL pagination method https://grey-panther.net/2008/10/efficient-sql-pagination-method.html https://grey-panther.net/2008/10/efficient-sql-pagination-method.html#comments Thu, 02 Oct 2008 05:34:00 +0000 https://grey-panther.net/?p=676 The usual technique for displaying data from an SQL query as multiple web-pages involves the LIMIT/OFFSET clause. For example for the first page the query would look like something like:

SELECT foo, bar, baz FROM ozz WHERE ... ORDER BY ... OFFSET 0 LIMIT 10

For the second page you would do:

SELECT foo, bar, baz FROM ozz WHERE ... ORDER BY ... OFFSET 10 LIMIT 10

And so on. There are two disadvantages to this method:

  • The smaller one: it can loose or duplicate data (from the user’s point of view, not from the DB’s point of view). This happens if data is inserted/deleted at a location which (given our ordering criteria) falls before the current offset. In this case, going a page forward can result in: (a) seeing some elements from the previous page or (b) skipping over some elements (they are not shown either in the current or the next page).

    These are usually small problems because (a) mostly data paginated is fairly static (so the probability of insertion/deletion occurring is fairly small) (b) it is tolerable by users (they shouldn’t use this method as a way to search the data, rather they should use the specialized search features) and (c) the likelihood of accessing a page decreases with their offset (users are more likely to access the first few pages than the last ones), which is great, since the likelihood of these errors manifesting themselves increases with the page offset (higher offset means a higher chance that the insertion/deletion will occur before it)

  • The second problem, which has the potential of becoming a big one given enough database entries, is the one of performance. Even if you specify a LIMIT clause, the database internally still has to fetch OFFSET + LIMIT rows. This value can become quite large.

    Here the counterarguments are: most databases are quite small. These types of queries aren’t a problem for well speced out servers for tens of thousands of rows. Also, as mentioned before, the likelihood of an user accessing a high-number page (thus creating a high-offset query) is fairly small.

While there are some good arguments for the fact that this isn’t a big problem, I would still like to present an alternative solution which can be applied in the few cases when this is a problem. Of course, the solution itself creates and other set of problems which I’ll discuss.

The presumption behind this idea is that our ordering criteria is backed by an index (there is an index which it can rely on to do the ordering). If this isn’t the case, there are bigger problems. If this is the case, then we can use this index to fetch only the relevant data from the table.

The B-Tree indexes which you’ll find in most current RDBMS’s, support the “<” (less than) and “>” (greater than) operators. So, if we could translate the query “elements from offset 10” into “elements greater than X”, we could use the index and thus avoiding fetching more rows than needed.

We can do this by transmitting the relevant field of the last element in the get query instead of the page number. IE: instead of “…/foo?page=10” the URL would look like “…/foo?after=1234”. One word of caution: if we don’t want to expose the ordering criteria, this solution isn’t appropriate or symmetrical crypto needs to be used to protect the value.

Now we can rewrite our initial queries into:

SELECT foo, bar, baz, zed FROM ozz WHERE ... AND zed > 1234 ORDER BY zed ASC LIMIT N

But what about stepping backwards (clicking on the “previous page” link)? In this case we are interested in the biggest N elements (where N is the number of elements per page you want to display) smaller than the first element displayed on the page. Again, we transmit this using the get query, writing instead of “…/foo?page=9” the URL “…/foo?before=1200” (supposing that we clicked “previous page” from page 10 – if we would have clicked “next page” from page 8, we would still have an “after” query). Now the query becomes:

SELECT foo, bar, baz, zed FROM ozz WHERE ... AND zed < 1200 ORDER BY zed DESC LIMIT N

One observation: the results returned by this query are in reverse of the order you would like to display them in, so you need to reverse them in code before displaying. If you don't want to change your code at all, you can use the DB to do this for you like this (this will be still fairly efficient, since the final sorting is applied only to the retrieved subset):

SELECT * FROM 
  (SELECT foo, bar, baz, zed 
    FROM ozz 
    WHERE ... AND zed < 1200 
    ORDER BY zed DESC LIMIT N)
  ORDER BY zed ASC

Two observations:

The column we base our sorting criteria on got added in the rows we retrieve. This is necessary because we need to transmit the value of this column for the first and and last element from the page through the "previous" and "next" links.

Also, in all the examples I supposed that the sorting order we want to display the elements in on the page is ascending. If it is descending, things must be swapped around.

Finally, what are the advantages / disadvantage of this method?

  • It solves the problem of loosing/duplicating elements when they get inserted/deleted before the current elements.
  • On the flipside, it can't guarantee you that (a) when you go next -> next -> next and then previous -> previous -> previous you will get the same elements on the page (if insertions / deletions occurred) and even worse (b) you can end up with fewer than N elements on the first page (which might be surprising for your visitors. This can be solved by handling the first page in a "special" way.
  • It is more complex to implement and might not fit well with existing query abstraction solutions (ORMs for example)
  • There isn't really an easy way to jump X pages forward or backward or to estimate the "number of pages". It might be acceptable for you to provide only "forward" and "back" links, without direct links to specific pages. Going X pages in either direction can be implemented by repeatedly issuing the query. To estimate the number of pages, you might do a "SELECT COUNT(1) FROM ... WHERE ..." for your criteria, which might be more memory efficient than actually fetching all the rows (depending on your RDBMS)

As always, there is no clear cut "best" solution, you must look at the restrictions your particular project has and decide which solution applies best to it.

Update: See an other take on this matter on the MySQL performance blog - Four ways to optimize paginated displays.

]]>
https://grey-panther.net/2008/10/efficient-sql-pagination-method.html/feed 3 676