gpt4 book ai didi

SQL:计数和子查询

转载 作者:行者123 更新时间:2023-12-03 01:37:39 24 4
gpt4 key购买 nike

再次使用 count 和 sql

在 sqlite 上,我有表格

  • 论文:paper_id、doi、年份
  • 作者:paper_id、author_id、inst_id
  • 作者:author_id、姓名、名字
  • inst:inst_id、名称、see_id

inst 是一个机构表:大学等。writeby 中的每一行给出一篇论文、一位作者、该作者当时所属的机构。可以有多个机构,并且每个机构都会重复一对 paper_id、author_id。对于给定的作者,我想要一个包含 paper.doi、papers.year 的列表以及与他合作撰写论文的合作者数量。我试过了

 SELECT  papers.doi, papers.year, count(*) as c
FROM authors
INNER JOIN writtenby ON authors.author_id = writtenby.author_id
INNER JOIN writtenby AS writtenby_1 ON writtenby.paper_id =
writtenby_1.paper_id
INNER JOIN papers on writtenby_1.paper_id = papers.paper_id
WHERE authors.name ='Beck' AND authors.firstname= 'H P'
GROUP BY papers.doi, papers.year
ORDER BY c DESC

我遇到的问题可能是,如果我正在搜索的作者在给定论文中出现两次(因为有两个机构)计数加倍。对于给定的论文,预期结果为 2890,由行数给出

SELECT DISTINCT author_id
FROM writtenby
WHERE paper_id = 4593

(我的数据:2890 行)如果没有 unique,我将有 3023 行,上面的第一个查询给出的计数为 6046。我尝试在上面的 Count 子句中使用 DISTINCT,但这仍然不起作用。

我可以在子查询中使用 count 吗?感谢您的帮助...

示例数据:

-- Make the tables

CREATE TABLE 'authors' (name collate nocase, firstname collate nocase, see_id integer, 'author_id' INTEGER PRIMARY KEY NOT NULL );
CREATE TABLE 'inst' ('name' TEXT NOT NULL, 'country' TEXT NOT NULL , 'see_id' INTEGER, 'inst_id' INTEGER PRIMARY KEY NOT NULL );
CREATE TABLE 'papers' ('doi' TEXT NOT NULL,'year' TEXT NOT NULL, 'paper_id' INTEGER PRIMARY KEY NOT NULL );
CREATE TABLE 'writtenby' ('paper_id' INTEGER NOT NULL, 'author_id' INTEGER NOT NULL, 'inst_id' INTEGER NOT NULL, PRIMARY KEY ('paper_id', 'author_id', 'inst_id'));

-- Insert the data

-- authors : 5 names, one with 2 variants

INSERT INTO 'authors' (name, firstname, see_id, author_id) VALUES ('Doe', 'J', 1, 1);
INSERT INTO 'authors' (name, firstname, see_id, author_id) VALUES ('Klein', 'K', 2, 2);
INSERT INTO 'authors' (name, firstname, see_id, author_id) VALUES ('Lang', 'F', 3, 3);
INSERT INTO 'authors' (name, firstname, see_id, author_id) VALUES ('Rue', 'A De La', 6, 4);
INSERT INTO 'authors' (name, firstname, see_id, author_id) VALUES ('La Rue', 'A De', 6, 5);
INSERT INTO 'authors' (name, firstname, see_id, author_id) VALUES ('De La Rue', 'A', 6, 6);
INSERT INTO 'authors' (name, firstname, see_id, author_id) VALUES ('Smith', 'S', 7, 7);

-- inst 4 name, 2 variants

INSERT INTO 'inst' (name, country, see_id, inst_id) VALUES ('Universite de Paris', 'France', 1, 1);
INSERT INTO 'inst' (name, country, see_id, inst_id) VALUES ('Paris University', 'France', 1, 2);
INSERT INTO 'inst' (name, country, see_id, inst_id) VALUES ('Universite de Lyon', 'France', 3, 3);
INSERT INTO 'inst' (name, country, see_id, inst_id) VALUES ('Univ Freiburg', 'Germany', 4, 4);
INSERT INTO 'inst' (name, country, see_id, inst_id) VALUES ('EPFZ', 'Switzerland', 5, 5);
INSERT INTO 'inst' (name, country, see_id, inst_id) VALUES ('Eidg Techn Hochschule', 'Switzerland', 5, 6);

-- papers: 3 papers

INSERT INTO 'papers' (doi, year, paper_id) VALUES ('doi1', '2017', 1);
INSERT INTO 'papers' (doi, year, paper_id) VALUES ('doi2', '2018', 2);
INSERT INTO 'papers' (doi, year, paper_id) VALUES ('doi3', '2018', 3);

-- paper 1: 4 authors

INSERT INTO 'writtenby' (paper_id, author_id, inst_id) VALUES (1, 6, 1);
INSERT INTO 'writtenby' (paper_id, author_id, inst_id) VALUES (1, 6, 3);
INSERT INTO 'writtenby' (paper_id, author_id, inst_id) VALUES (1, 1, 5);
INSERT INTO 'writtenby' (paper_id, author_id, inst_id) VALUES (1, 2, 4);
INSERT INTO 'writtenby' (paper_id, author_id, inst_id) VALUES (1, 7, 1);

-- paper 2: 3 authors

INSERT INTO 'writtenby' (paper_id, author_id, inst_id) VALUES (2, 6, 1);
INSERT INTO 'writtenby' (paper_id, author_id, inst_id) VALUES (2, 6, 3);
INSERT INTO 'writtenby' (paper_id, author_id, inst_id) VALUES (2, 1, 5);
INSERT INTO 'writtenby' (paper_id, author_id, inst_id) VALUES (2, 2, 5);

-- paper 3: 3 authors

INSERT INTO 'writtenby' (paper_id, author_id, inst_id) VALUES (3, 6, 1);
INSERT INTO 'writtenby' (paper_id, author_id, inst_id) VALUES (3, 2, 4);
INSERT INTO 'writtenby' (paper_id, author_id, inst_id) VALUES (3, 6, 3);
INSERT INTO 'writtenby' (paper_id, author_id, inst_id) VALUES (3, 2, 1);
INSERT INTO 'writtenby' (paper_id, author_id, inst_id) VALUES (3, 3, 4);
INSERT INTO 'writtenby' (paper_id, author_id, inst_id) VALUES (3, 3, 5);
INSERT INTO 'writtenby' (paper_id, author_id, inst_id) VALUES (3, 3, 1);

检查查询:

 SELECT  papers.doi, papers.year, count(*) as c
FROM authors
INNER JOIN writtenby ON authors.author_id = writtenby.author_id
INNER JOIN writtenby AS writtenby_1 ON writtenby.paper_id =
writtenby_1.paper_id
INNER JOIN papers on writtenby_1.paper_id = papers.paper_id
WHERE authors.name ='De La Rue' AND authors.firstname= 'A'
GROUP BY papers.doi, papers.year
ORDER BY c DESC


SELECT p.doi, p.year, COUNT(w2.author_id) AS cnt
FROM authors a
INNER JOIN writtenby w1
ON a.author_id = w1.author_id
INNER JOIN writtenby w2
ON w1.paper_id = w2.paper_id AND w1.author_id <> w2.author_id
INNER JOIN papers p
ON w2.paper_id = p.paper_id
WHERE
a.name = 'De La Rue' AND a.firstname = 'A'
GROUP BY
p.doi, p.year
ORDER BY
cnt DESC;

两个查询都给出了错误的结果第一个:

doi3|2018|14
doi1|2017|10
doi2|2018|8

第二个查询

doi3|2018|10
doi1|2017|6
doi2|2018|4

弗朗索瓦

最佳答案

我发现正在发生的一个计数问题是在 writingby 表的自联接中。在那里,您不会检查匹配行是否具有不同 author_id。如果 author_id 相同,那么您不应该计算它。此外,您应该计算第二个 writingby 表的共享作者数量。这样,如果给定作者没有任何共同作者,计数将显示为零。

SELECT p.doi, p.year, COUNT(w2.author_id) AS cnt
FROM authors a
INNER JOIN writtenby w1
ON a.author_id = w1.author_id
INNER JOIN writtenby w2
ON w1.paper_id = w2.paper_id AND w1.author_id <> w2.author_id
INNER JOIN papers p
ON w2.paper_id = p.paper_id
WHERE
a.name = 'Beck' AND a.firstname = 'H P'
GROUP BY
p.doi, p.year
ORDER BY
cnt DESC;

关于SQL:计数和子查询,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/54111382/

24 4 0
Copyright 2021 - 2024 cfsdn All Rights Reserved 蜀ICP备2022000587号
广告合作:1813099741@qq.com 6ren.com