Q200510-02-02: 重复的DNA序列 SQL解法

重复的DNA序列
所有 DNA 都由一系列缩写为 A,C,G 和 T 的核苷酸组成,例如:“ACGAATTCCG”。在研究 DNA 时,识别 DNA 中的重复序列有时会对研究非常有帮助。

编写一个函数来查找 DNA 分子中所有出现超过一次的 10 个字母长的序列(子串)。

示例:

输入:s = "AAAAACCCCCAAAAACCCCCCAAAAAGGGTTT"
输出:["AAAAACCCCC", "CCCCCAAAAA"]

 

使用Oracle11g数据库

最终SQL及结果:

复制代码
SQL> select sub from
  2  (select count(*) as cnt,sub from
  3  (select substr(a.sery,b.rn,10) as sub
  4  from tb_string a,
  5  (select rownum as rn from dual connect by level<=(select length(max(sery)) from tb_string )) b) c
  6  group by sub) d
  7  where d.cnt>1;

SUB
--------------------
AAAAACCCCC
CCCCCAAAAA
复制代码

思考过程:

复制代码
create table tb_string(
    id number(4,0) primary key,
    sery nvarchar2(50) not null)
    
insert into tb_string(id,sery) values('1','AAAAACCCCCAAAAACCCCCCAAAAAGGGTTT');

select rownum
from dual
connect by level<=(select length(max(sery)) from tb_string );

select *
from tb_string,
(select rownum from dual connect by level<=(select length(max(sery)) from tb_string )) a

select substr(a.sery,b.rn,10) as sub
from tb_string a,
(select rownum as rn from dual connect by level<=(select length(max(sery)) from tb_string )) b

select count(*) as cnt,sub from
(select substr(a.sery,b.rn,10) as sub
from tb_string a,
(select rownum as rn from dual connect by level<=(select length(max(sery)) from tb_string )) b) c
group by sub

select sub from
(select count(*) as cnt,sub from
(select substr(a.sery,b.rn,10) as sub
from tb_string a,
(select rownum as rn from dual connect by level<=(select length(max(sery)) from tb_string )) b) c
group by sub) d
where d.cnt>1
复制代码

--2020年5月11日 --

 

posted @   逆火狂飙  阅读(140)  评论(1编辑  收藏  举报
编辑推荐:
· Linux系列:如何用heaptrack跟踪.NET程序的非托管内存泄露
· 开发者必知的日志记录最佳实践
· SQL Server 2025 AI相关能力初探
· Linux系列:如何用 C#调用 C方法造成内存泄露
· AI与.NET技术实操系列(二):开始使用ML.NET
阅读排行:
· 无需6万激活码!GitHub神秘组织3小时极速复刻Manus,手把手教你使用OpenManus搭建本
· C#/.NET/.NET Core优秀项目和框架2025年2月简报
· Manus爆火,是硬核还是营销?
· 终于写完轮子一部分:tcp代理 了,记录一下
· 【杭电多校比赛记录】2025“钉耙编程”中国大学生算法设计春季联赛(1)
生当作人杰 死亦为鬼雄 至今思项羽 不肯过江东
点击右上角即可分享
微信分享提示