ansaurus

Question

Cycle detection with recursive subquery factoring

Answer 1

+1 A:

MySQL Server version 5.0.45 didn't like with:
ERROR 1064 (42000): You have an error in your SQL syntax; check the manual that corresponds to your MySQL server version for the right syntax to use near 'with tr (id, parent_id) as (select id, parent_id from t where id = 1 union all s' at line 1

wallyk 2009-11-13 21:08:50

Thanks for trying, wallyk. But does MySQL support something similar and this just isn't the right syntax, or doesn't MySQL doesn't support recursive queries at all?

Rob van Wijk 2009-11-13 21:51:02

I don't think it does. It was only a few years ago that it did not support stored procedures, so it has evolved rather quickly. I wouldn't be surprised if such things are in the latest version, or a upcoming version.

wallyk 2009-11-13 22:02:19

Answer 2

+2 A:

AFAIK:

MySQL doesn't support recursive CTE's
SQL Sever does not support cycle detection in recursive CTE's

Andomar 2009-11-13 23:03:43

MySQL doesn't support CTEs at all.

OMG Ponies 2009-11-13 23:07:10

That's clear, thanks. So MySQL and SQL Server cannot be used to test this script.

Rob van Wijk 2009-11-14 09:25:06

Indeed, MS SQL Server has a maximum recursion limit, something like 100 by default.

Vilx- 2009-11-16 09:43:11

Answer 3

+2 A:

PostgreSQL supports WITH-style hierarchical queries, but doesn't have any automatic cycle detection. This means that you need to write your own and the number of rows returned depends on the way you specify join conditions in the recursive part of the query.

Both examples use an array if IDs (called all_ids) to detect loops:

WITH recursive tr (id, parent_id, all_ids, cycle) AS (
    SELECT id, parent_id, ARRAY[id], false
    FROM t
    WHERE id = 1
    UNION ALL
    SELECT t.id, t.parent_id, all_ids || t.id, t.id = ANY(all_ids)
    FROM t
    JOIN tr ON t.parent_id = tr.id AND NOT cycle)
SELECT id, parent_id, cycle
FROM tr;

 id | parent_id | cycle
----+-----------+-------
  1 |         2 | f
  2 |         1 | f
  1 |         2 | t


WITH recursive tr (id, parent_id, all_ids, cycle) AS (
    SELECT id, parent_id, ARRAY[id], false
    FROM t
    WHERE id = 1
    UNION ALL
    SELECT t.id, t.parent_id, all_ids || t.id, (EXISTS(SELECT 1 FROM t AS x WHERE x.id = t.parent_id))
    FROM t
    JOIN tr ON t.parent_id = tr.id
    WHERE NOT t.id = ANY(all_ids))
SELECT id, parent_id, cycle
FROM tr;

 id | parent_id | cycle
----+-----------+-------
  1 |         2 | f
  2 |         1 | t

Aleksander Kmetec 2009-11-14 01:06:56

Very interesting queries. Oracle doesn't support that array syntax, but it's nice to see that the results are the same as in Oracle's ANSI syntax. It leads me to think that the ANSI query maybe doesn't contain a bug after all. Thanks, Aleksander.

Rob van Wijk 2009-11-14 09:32:59

Answer 4

+3 A:

From documentation on CONNECT_BY_ISCYCLE:

The CONNECT_BY_ISCYCLE pseudocolumn returns 1 if the current row has a child which is also its ancestor

and that on CYCLE:

A row is considered to form a cycle if one of its ancestor rows has the same values for the cycle columns.

In your example, row 2 does have a child which is also its ancestor, but its id has not been returned yet.

In other words, CONNECT_BY_ISCYCLE checks the children (which are yet to be returned), while CYCLE checks the current row (which is already returned).

CONNECT BY is row based, while recursive CTE's are set-based.

There is no concept of a "child" in a recursive CTE. It's a set based operation which can yield results completely out of the tree. Generally speaking, the anchor part and the recursive part can even use the different tables.

Since recursive CTE's are usually used to build hierarchy trees, Oracle decided to add a cycle check. But due the set-based way the recursive CTE's operate, it's generally impossible to tell will the next step generate a cycle or not.

To perform the "next" step, the whole "current" set needs to be available, but to generate each row of the current set (which includes the cycle column) we just need to have the results of the "next" operation. It's not a problem with a single row (like in CONNECT BY), but it is a problem with a set as a whole.

Didn't look into Oracle 11 yet, but SQL Server implements recursive CTE's by just hiding a CONNECT BY behind them, which requires placing numerous restrictions (all of which effectively forbid all set-based operations).

PostgreSQL's implementation, on the other hand, is truly set-based.

As was mentioned before, MySQL does not implement CTE's at all (it does not implement HASH JOIN's or MERGE JOINs as well, only the nested loops, so don't be surprised much).

Ironically, I received a letter today on this very subject, which I will cover in my blog.

Update:

Recursive CTE's in SQL Server are no more than CONNECT BY in disguise. See this article in my blog for shocking details:

SQL Server: are the recursive CTE’s really set-based?

Quassnoi 2009-11-18 18:02:20

Good explanation, Quassnoi. Thanks. So the recursive with clause works as expected. It's just different than the connect-by. It's nice to be aware of these little nuances between the two.

Rob van Wijk 2009-11-18 22:42:04

ansaurus

tags:

views:

answers:

Cycle detection with recursive subquery factoring

related questions